AE Agent Errors
AE-6QKQR78H

ReAct agent emits a tool Action but never runs the tool, fabricating the Observation and Final Answer itself

HALLUCINATIONseverity: HIGHcause: LIKELYoutcome: MITIGATEDconfidence: LOW

CrewAI agents produced a valid-looking Thought/Action/Observation/Final Answer trace while the tool was never invoked: the model wrote the Observation itself and the framework accepted the Final Answer. Reports span custom BaseTool tools, MCP tools and standard tools, and several users observed it with GPT-5/o-series models but not GPT-4.1.

Framework / agent
CrewAI · CrewAI tool-using agent (text ReAct loop)
Remediation attempts
FAILEDPARTIAL SUCCESSSUGGESTEDSUGGESTEDSUGGESTED
Recurrence
not documented
Source languages
en
Updated
2026-09-30

Sources

Public sample: the full evidence record is free.

Symptoms

  • The agent's trace shows an Action and an Observation, but the tool's run() is never executed and no tool activity appears in traces REPORTED CLAIM
  • The Observation content is generated by the model and the agent finishes with a Final Answer built on it REPORTED CLAIM
  • Fabrication appears after switching to GPT-5 while GPT-4.1 executes tools correctly REPORTED CLAIM

Context and trigger

Agent type
CrewAI tool-using agent (text ReAct loop)
Component
text-based ReAct parsing in the agent executor (process_llm_response / CrewAgentParser)
Framework
CrewAI (0.141.0 (report); main 1.15.2a2 (root-cause analysis))
Model
Qwen2.5-72B-Instruct-GPTQ-Int4 (report); GPT-5 in later reports
Task
research / file tasks using web search, MCP and custom tools
Tools
custom BaseTool web search, MCP tools over SSE, Serper search
  • The model writes the Observation and Final Answer in the same completion as the Action REPORTED CLAIM
  • The model does not support stop words (e.g. GPT-5/o-series), so generation is not cut at the Observation REPORTED CLAIM

Open questions

  • The maintainers did not confirm the root cause; both the issue and the root-cause issue were closed as not planned.
  • Whether the reporter's later successful test used the maintainer branch or a third-party verification PR is ambiguous in the thread.

Root cause LIKELY

For models without stop-word support the completion runs past the real Action and includes a fabricated Observation and Final Answer; the parser returns AgentFinish whenever 'Final Answer:' appears, and the recovery block meant to truncate at the Observation has been unreachable since the parser error it relies on was removed.

cause reported, not confirmed by the project

the LLM generates straight past the real tool call and fabricates the rest of the ReAct loop in a single completionreporter · github.com/crewAIInc/crewAI
Since then `parse()` returns the fabricated `AgentFinish` whenever `Final Answer:` appears anywhere in the text, and the recovery can never run.reporter · github.com/crewAIInc/crewAI

Remediation attempts (5)

FAILEDparser_fixby contributor

Maintainer-team work-in-progress branch reworking tool invocation; the reporter tested it and saw no change.

Status basis: documented as not fixing the failure. Verification: not verified.

I tried running 3 tests with this version, but it still behaves the same as beforereporter · github.com/crewAIInc/crewAI
PARTIAL SUCCESSprompt_changeby third_party

Workaround: modify the tool-usage prompt (en.json) so the model does not write the Observation itself. Some users report no failures in dozens of attempts, others still see fabrication or loops.

Status basis: documented as only partly fixing the failure. Verification: not verified.

I tried modifying the en.json file as you suggested, but the results are still inconsistent.third_party · github.com/crewAIInc/crewAI
I've tested it with both gpt-4.1-mini, and gpt-5-mini, no failures yet on a few dozen attempts.third_party · github.com/crewAIInc/crewAI
SUGGESTEDworkaroundby third_party

Workaround: use GPT-4.1 instead of GPT-5/o-series models.

Status basis: proposed; no evidence it was applied. Verification: not verified.

the issue was fixed for me when I used GPT 4.1 instead of the 'o' reasoning models from OpenAI.third_party · github.com/crewAIInc/crewAI
SUGGESTEDparser_fixby third_party

Replace the dead recovery block with an explicit position check that truncates at the fabricated Observation so the real Action executes, with fails-before/passes-after tests. Independently re-verified by a third party but closed without merge.

Status basis: proposed; no evidence it was applied. Verification: not verified.

Change: https://github.com/crewAIInc/crewAI/pull/6450 · not merged · tests changed

Excerpt of the change (lib/crewai/src/crewai/utilities/agent_utils.py, MIT):

@@ -17,7 +17,7 @@
 from pydantic import BaseModel
 from rich.console import Console
 
-from crewai.agents.constants import FINAL_ANSWER_AND_PARSABLE_ACTION_ERROR_MESSAGE
+from crewai.agents.constants import ACTION_INPUT_REGEX, FINAL_ANSWER_ACTION
 from crewai.agents.parser import (
     AgentAction,
     AgentFinish,
@@ -576,12 +576,20 @@ def process_llm_response(
     Returns:
         Either an AgentAction or AgentFinish
     """
-    if not use_stop_words:
-        try:
-            format_answer(answer)
-        except OutputParserError as e:
-            if FINAL_ANSWER_AND_PARSABLE_ACTION_ERROR_MESSAGE in e.error:
-                answer = answer.split("Observation:")[0].strip()
+    if not use_stop_words and FINAL_ANSWER_ACTION in answer:
+        action_match = ACTION_INPUT_REGEX.search(answer)
+        final_answer_idx = answer.find(FINAL_ANSWER_ACTION)
+        if action_match and action_match.start() < final_answer_idx:
+            # Without the "\nObservation:" stop sequence the model generates past
+            # the real tool call, fabricating an Observation and Final Answer.
+            # Discard the fabricated continuation so the actual Action executes.
+            # Anchor on the newline (the real stop sequence) so an "Observation:"
+            # substring inside the Action Input payload isn't mistaken for it.
+            observation_idx = answer.find(
+                "\nObservation:", action_match.start(), final_answer_idx
Replace the dead `except` block with an explicit position checkthird_party · github.com/crewAIInc/crewAI
applying this PR's diff locally, the same input returns `AgentAction(tool="web_search", ...)` with the fabricated `Observation:`/`Final Answer` truncatedthird_party · github.com/crewAIInc/crewAI
SUGGESTEDotherby third_party

Third-party tool-execution authenticity verification system (filesystem/subprocess monitoring); closed without merge.

Status basis: proposed; no evidence it was applied. Verification: not verified.

Change: https://github.com/crewAIInc/crewAI/pull/3378 · not merged

Excerpt of the change (demo_tool_verification.py, MIT):

@@ -0,0 +1,139 @@
+#!/usr/bin/env python3
+"""
+Tool Execution Verification Demo
+
+This script demonstrates the tool execution verification system by testing
+real vs fake tool implementations. It shows how the system can detect when
+tools are actually executing vs when they're fabricating results.
+
+Usage:
+    python demo_tool_verification.py
+
+The demo will:
+1. Test a real file writing tool that actually creates files
+2. Test a fake file writing tool that only pretends to create files
+3. Show the verification results for each
+"""
+
+import os
+import sys
+import tempfile
+from pathlib import Path
+
+# Add the src directory to the path so we can import our modules
+sys.path.insert(0, str(Path(__file__).parent / "src"))
+
+from crewai.utilities.tool_execution_verifier import (
+    verify_tool_execution
+)
+
Added real-time tool execution monitoring system that:contributor · github.com/crewAIInc/crewAI

Outcome outcome: MITIGATED

only partial remediation documented

Recurrence

Not documented in the sources (absence of reports is not evidence of absence).

Confidence LOW

Cause or remediation is reported, but not confirmed by the affected project.

FactorPresentMeaning
first_party_evidencenoa quote from the affected project/vendor (or a controlled test)
fix_appliednoa fix was merged/released
regression_testyestests changed with the fix
independent_confirmationnoreporter/maintainer/vendor confirmed the failure is gone
root_cause_verifiednocause stated by the project and addressed by the fix
reproduction_documentednosteps or conditions to reproduce were quoted
multiple_independent_sourcesno1 independent source group(s)
failed_attempts_documentedyesunsuccessful remediation recorded

All evidence (12 verified quotes)

the agent **does not actually invoke the tool at runtime**, even though it produces a valid-looking `Thought → Action → Observation → Final Answer` trace.reporter · github.com/crewAIInc/crewAI
Instead of executing the tool (e.g., calling `tool.run()`), the LLM **generates a fake Observation output** and continues to the final answer.reporter · github.com/crewAIInc/crewAI
However, when switching to GPT-5, the agent generates a completely fabricated Observation.third_party · github.com/crewAIInc/crewAI
I tried running 3 tests with this version, but it still behaves the same as beforereporter · github.com/crewAIInc/crewAI
I tried modifying the en.json file as you suggested, but the results are still inconsistent.third_party · github.com/crewAIInc/crewAI
I've tested it with both gpt-4.1-mini, and gpt-5-mini, no failures yet on a few dozen attempts.third_party · github.com/crewAIInc/crewAI
the LLM generates straight past the real tool call and fabricates the rest of the ReAct loop in a single completionreporter · github.com/crewAIInc/crewAI
Since then `parse()` returns the fabricated `AgentFinish` whenever `Final Answer:` appears anywhere in the text, and the recovery can never run.reporter · github.com/crewAIInc/crewAI
Replace the dead `except` block with an explicit position checkthird_party · github.com/crewAIInc/crewAI
applying this PR's diff locally, the same input returns `AgentAction(tool="web_search", ...)` with the fabricated `Observation:`/`Final Answer` truncatedthird_party · github.com/crewAIInc/crewAI
the issue was fixed for me when I used GPT 4.1 instead of the 'o' reasoning models from OpenAI.third_party · github.com/crewAIInc/crewAI
Added real-time tool execution monitoring system that:contributor · github.com/crewAIInc/crewAI

Categories: HALLUCINATION TOOL FAILURE · extraction curated, gate-2026-09-29-v1

Similarity to your system is not implied. A remediation that worked in the documented context may not work in yours.