ReAct agent emits a tool Action but never runs the tool, fabricating the Observation and Final Answer itself
CrewAI agents produced a valid-looking Thought/Action/Observation/Final Answer trace while the tool was never invoked: the model wrote the Observation itself and the framework accepted the Final Answer. Reports span custom BaseTool tools, MCP tools and standard tools, and several users observed it with GPT-5/o-series models but not GPT-4.1.
- Framework / agent
- CrewAI · CrewAI tool-using agent (text ReAct loop)
- Remediation attempts
- FAILEDPARTIAL SUCCESSSUGGESTEDSUGGESTEDSUGGESTED
- Recurrence
- not documented
- Source languages
- en
- Updated
- 2026-09-30
Sources
- GITHUB ISSUE [BUG] 🐞Agent does not actually invoke tools, only simulates tool usage with fabricated output — github.com/crewAIInc/crewAI, retrieved 2026-09-29
- GITHUB ISSUE [BUG] Fabricated-Observation recovery in process_llm_response is dead code since #2483 — real tool calls silently discarded for models without stop-word support — github.com/crewAIInc/crewAI, retrieved 2026-09-29
- GITHUB PULL REQUEST fix(agents): recover real tool call from fabricated Observation continuations — github.com/crewAIInc/crewAI, retrieved 2026-09-29
- GITHUB PULL REQUEST feat: add tool execution authenticity verification system (#3154) — github.com/crewAIInc/crewAI, retrieved 2026-09-29
Public sample: the full evidence record is free.
Symptoms
- The agent's trace shows an Action and an Observation, but the tool's run() is never executed and no tool activity appears in traces REPORTED CLAIM
- The Observation content is generated by the model and the agent finishes with a Final Answer built on it REPORTED CLAIM
- Fabrication appears after switching to GPT-5 while GPT-4.1 executes tools correctly REPORTED CLAIM
Context and trigger
- Agent type
- CrewAI tool-using agent (text ReAct loop)
- Component
- text-based ReAct parsing in the agent executor (process_llm_response / CrewAgentParser)
- Framework
- CrewAI (0.141.0 (report); main 1.15.2a2 (root-cause analysis))
- Model
- Qwen2.5-72B-Instruct-GPTQ-Int4 (report); GPT-5 in later reports
- Task
- research / file tasks using web search, MCP and custom tools
- Tools
- custom BaseTool web search, MCP tools over SSE, Serper search
- The model writes the Observation and Final Answer in the same completion as the Action REPORTED CLAIM
- The model does not support stop words (e.g. GPT-5/o-series), so generation is not cut at the Observation REPORTED CLAIM
Open questions
- The maintainers did not confirm the root cause; both the issue and the root-cause issue were closed as not planned.
- Whether the reporter's later successful test used the maintainer branch or a third-party verification PR is ambiguous in the thread.
Root cause LIKELY
For models without stop-word support the completion runs past the real Action and includes a fabricated Observation and Final Answer; the parser returns AgentFinish whenever 'Final Answer:' appears, and the recovery block meant to truncate at the Observation has been unreachable since the parser error it relies on was removed.
cause reported, not confirmed by the project
the LLM generates straight past the real tool call and fabricates the rest of the ReAct loop in a single completionSince then `parse()` returns the fabricated `AgentFinish` whenever `Final Answer:` appears anywhere in the text, and the recovery can never run.Remediation attempts (5)
Maintainer-team work-in-progress branch reworking tool invocation; the reporter tested it and saw no change.
Status basis: documented as not fixing the failure. Verification: not verified.
I tried running 3 tests with this version, but it still behaves the same as beforeWorkaround: modify the tool-usage prompt (en.json) so the model does not write the Observation itself. Some users report no failures in dozens of attempts, others still see fabrication or loops.
Status basis: documented as only partly fixing the failure. Verification: not verified.
I tried modifying the en.json file as you suggested, but the results are still inconsistent.I've tested it with both gpt-4.1-mini, and gpt-5-mini, no failures yet on a few dozen attempts.Workaround: use GPT-4.1 instead of GPT-5/o-series models.
Status basis: proposed; no evidence it was applied. Verification: not verified.
the issue was fixed for me when I used GPT 4.1 instead of the 'o' reasoning models from OpenAI.Replace the dead recovery block with an explicit position check that truncates at the fabricated Observation so the real Action executes, with fails-before/passes-after tests. Independently re-verified by a third party but closed without merge.
Status basis: proposed; no evidence it was applied. Verification: not verified.
Change: https://github.com/crewAIInc/crewAI/pull/6450 · not merged · tests changed
Excerpt of the change (lib/crewai/src/crewai/utilities/agent_utils.py, MIT):
@@ -17,7 +17,7 @@
from pydantic import BaseModel
from rich.console import Console
-from crewai.agents.constants import FINAL_ANSWER_AND_PARSABLE_ACTION_ERROR_MESSAGE
+from crewai.agents.constants import ACTION_INPUT_REGEX, FINAL_ANSWER_ACTION
from crewai.agents.parser import (
AgentAction,
AgentFinish,
@@ -576,12 +576,20 @@ def process_llm_response(
Returns:
Either an AgentAction or AgentFinish
"""
- if not use_stop_words:
- try:
- format_answer(answer)
- except OutputParserError as e:
- if FINAL_ANSWER_AND_PARSABLE_ACTION_ERROR_MESSAGE in e.error:
- answer = answer.split("Observation:")[0].strip()
+ if not use_stop_words and FINAL_ANSWER_ACTION in answer:
+ action_match = ACTION_INPUT_REGEX.search(answer)
+ final_answer_idx = answer.find(FINAL_ANSWER_ACTION)
+ if action_match and action_match.start() < final_answer_idx:
+ # Without the "\nObservation:" stop sequence the model generates past
+ # the real tool call, fabricating an Observation and Final Answer.
+ # Discard the fabricated continuation so the actual Action executes.
+ # Anchor on the newline (the real stop sequence) so an "Observation:"
+ # substring inside the Action Input payload isn't mistaken for it.
+ observation_idx = answer.find(
+ "\nObservation:", action_match.start(), final_answer_idxReplace the dead `except` block with an explicit position checkapplying this PR's diff locally, the same input returns `AgentAction(tool="web_search", ...)` with the fabricated `Observation:`/`Final Answer` truncatedThird-party tool-execution authenticity verification system (filesystem/subprocess monitoring); closed without merge.
Status basis: proposed; no evidence it was applied. Verification: not verified.
Change: https://github.com/crewAIInc/crewAI/pull/3378 · not merged
Excerpt of the change (demo_tool_verification.py, MIT):
@@ -0,0 +1,139 @@ +#!/usr/bin/env python3 +""" +Tool Execution Verification Demo + +This script demonstrates the tool execution verification system by testing +real vs fake tool implementations. It shows how the system can detect when +tools are actually executing vs when they're fabricating results. + +Usage: + python demo_tool_verification.py + +The demo will: +1. Test a real file writing tool that actually creates files +2. Test a fake file writing tool that only pretends to create files +3. Show the verification results for each +""" + +import os +import sys +import tempfile +from pathlib import Path + +# Add the src directory to the path so we can import our modules +sys.path.insert(0, str(Path(__file__).parent / "src")) + +from crewai.utilities.tool_execution_verifier import ( + verify_tool_execution +) +
Added real-time tool execution monitoring system that:
Outcome outcome: MITIGATED
only partial remediation documented
Recurrence
Not documented in the sources (absence of reports is not evidence of absence).
Confidence LOW
Cause or remediation is reported, but not confirmed by the affected project.
| Factor | Present | Meaning |
|---|---|---|
| first_party_evidence | no | a quote from the affected project/vendor (or a controlled test) |
| fix_applied | no | a fix was merged/released |
| regression_test | yes | tests changed with the fix |
| independent_confirmation | no | reporter/maintainer/vendor confirmed the failure is gone |
| root_cause_verified | no | cause stated by the project and addressed by the fix |
| reproduction_documented | no | steps or conditions to reproduce were quoted |
| multiple_independent_sources | no | 1 independent source group(s) |
| failed_attempts_documented | yes | unsuccessful remediation recorded |
All evidence (12 verified quotes)
the agent **does not actually invoke the tool at runtime**, even though it produces a valid-looking `Thought → Action → Observation → Final Answer` trace.Instead of executing the tool (e.g., calling `tool.run()`), the LLM **generates a fake Observation output** and continues to the final answer.However, when switching to GPT-5, the agent generates a completely fabricated Observation.I tried running 3 tests with this version, but it still behaves the same as beforeI tried modifying the en.json file as you suggested, but the results are still inconsistent.I've tested it with both gpt-4.1-mini, and gpt-5-mini, no failures yet on a few dozen attempts.the LLM generates straight past the real tool call and fabricates the rest of the ReAct loop in a single completionSince then `parse()` returns the fabricated `AgentFinish` whenever `Final Answer:` appears anywhere in the text, and the recovery can never run.Replace the dead `except` block with an explicit position checkapplying this PR's diff locally, the same input returns `AgentAction(tool="web_search", ...)` with the fabricated `Observation:`/`Final Answer` truncatedthe issue was fixed for me when I used GPT 4.1 instead of the 'o' reasoning models from OpenAI.Added real-time tool execution monitoring system that:
Categories: HALLUCINATION TOOL FAILURE · extraction curated, gate-2026-09-29-v1
Similarity to your system is not implied. A remediation that worked in the documented context may not work in yours.