{"schema_version":"1","case_id":"AE-6QKQR78H","case_type":"REAL_WORLD","failure":{"title":"ReAct agent emits a tool Action but never runs the tool, fabricating the Observation and Final Answer itself","category":"HALLUCINATION","secondary_categories":["TOOL_FAILURE"],"description":"CrewAI agents produced a valid-looking Thought/Action/Observation/Final Answer trace while the tool was never invoked: the model wrote the Observation itself and the framework accepted the Final Answer. Reports span custom BaseTool tools, MCP tools and standard tools, and several users observed it with GPT-5/o-series models but not GPT-4.1.","symptoms":[{"text":"The agent's trace shows an Action and an Observation, but the tool's run() is never executed and no tool activity appears in traces","basis":"REPORTED_CLAIM","evidence_ids":["e1","e2"]},{"text":"The Observation content is generated by the model and the agent finishes with a Final Answer built on it","basis":"REPORTED_CLAIM","evidence_ids":["e2","e3"]},{"text":"Fabrication appears after switching to GPT-5 while GPT-4.1 executes tools correctly","basis":"REPORTED_CLAIM","evidence_ids":["e3"]}],"error_messages":[],"severity":"HIGH","severity_basis":"Agent silently returns fabricated tool results as if the tool ran; users report it makes tool use unreliable, including for write operations.","affected_component":"text-based ReAct parsing in the agent executor (process_llm_response / CrewAgentParser)","agent_type":"CrewAI tool-using agent (text ReAct loop)","framework":{"name":"CrewAI","versions":"0.141.0 (report); main 1.15.2a2 (root-cause analysis)"},"model":"Qwen2.5-72B-Instruct-GPTQ-Int4 (report); GPT-5 in later reports","environment":{"language":"python","os":"Ubuntu 22.04"}},"trigger":{"conditions":[{"text":"The model writes the Observation and Final Answer in the same completion as the Action","basis":"REPORTED_CLAIM","evidence_ids":["e7"]},{"text":"The model does not support stop words (e.g. GPT-5/o-series), so generation is not cut at the Observation","basis":"REPORTED_CLAIM","evidence_ids":["e7","e3"]}],"task_type":"research / file tasks using web search, MCP and custom tools","tool_context":["custom BaseTool web search","MCP tools over SSE","Serper search"],"input_pattern":null,"unknowns":["The maintainers did not confirm the root cause; both the issue and the root-cause issue were closed as not planned.","Whether the reporter's later successful test used the maintainer branch or a third-party verification PR is ambiguous in the thread."]},"root_cause":{"status":"LIKELY","claimed_status":"LIKELY","status_reason":"cause reported, not confirmed by the project","description":"For models without stop-word support the completion runs past the real Action and includes a fabricated Observation and Final Answer; the parser returns AgentFinish whenever 'Final Answer:' appears, and the recovery block meant to truncate at the Observation has been unreachable since the parser error it relies on was removed.","basis":"REPORTED_CLAIM","evidence_ids":["e7","e8"]},"remediation_attempts":[{"attempt_id":"r1","description":"Maintainer-team work-in-progress branch reworking tool invocation; the reporter tested it and saw no change.","method":"parser_fix","actor":"contributor","status":"FAILED","claimed_status":"FAILED","status_reason":"documented as not fixing the failure","code_or_patch_reference":null,"released_in":null,"verification":{"method":"reporter_confirmation","evidence_ids":[],"details":"not verified"},"evidence_ids":["e4"]},{"attempt_id":"r2","description":"Workaround: modify the tool-usage prompt (en.json) so the model does not write the Observation itself. Some users report no failures in dozens of attempts, others still see fabrication or loops.","method":"prompt_change","actor":"third_party","status":"PARTIAL_SUCCESS","claimed_status":"PARTIAL_SUCCESS","status_reason":"documented as only partly fixing the failure","code_or_patch_reference":null,"released_in":null,"verification":{"method":"none","evidence_ids":[],"details":"not verified"},"evidence_ids":["e5","e6"]},{"attempt_id":"r3","description":"Workaround: use GPT-4.1 instead of GPT-5/o-series models.","method":"workaround","actor":"third_party","status":"SUGGESTED","claimed_status":"SUGGESTED","status_reason":"proposed; no evidence it was applied","code_or_patch_reference":null,"released_in":null,"verification":{"method":"none","evidence_ids":[],"details":"not verified"},"evidence_ids":["e11"]},{"attempt_id":"r4","description":"Replace the dead recovery block with an explicit position check that truncates at the fabricated Observation so the real Action executes, with fails-before/passes-after tests. Independently re-verified by a third party but closed without merge.","method":"parser_fix","actor":"third_party","status":"SUGGESTED","claimed_status":"SUGGESTED","status_reason":"proposed; no evidence it was applied","code_or_patch_reference":{"source_id":"gh:crewAIInc/crewAI#6450","url":"https://github.com/crewAIInc/crewAI/pull/6450","merged":false,"merged_at":null,"files":[{"path":"lib/crewai/src/crewai/utilities/agent_utils.py","is_test":false},{"path":"lib/crewai/tests/utilities/test_agent_utils.py","is_test":true}],"tests_changed":true,"snippet":{"path":"lib/crewai/src/crewai/utilities/agent_utils.py","language":"python","code":"@@ -17,7 +17,7 @@\n from pydantic import BaseModel\n from rich.console import Console\n \n-from crewai.agents.constants import FINAL_ANSWER_AND_PARSABLE_ACTION_ERROR_MESSAGE\n+from crewai.agents.constants import ACTION_INPUT_REGEX, FINAL_ANSWER_ACTION\n from crewai.agents.parser import (\n     AgentAction,\n     AgentFinish,\n@@ -576,12 +576,20 @@ def process_llm_response(\n     Returns:\n         Either an AgentAction or AgentFinish\n     \"\"\"\n-    if not use_stop_words:\n-        try:\n-            format_answer(answer)\n-        except OutputParserError as e:\n-            if FINAL_ANSWER_AND_PARSABLE_ACTION_ERROR_MESSAGE in e.error:\n-                answer = answer.split(\"Observation:\")[0].strip()\n+    if not use_stop_words and FINAL_ANSWER_ACTION in answer:\n+        action_match = ACTION_INPUT_REGEX.search(answer)\n+        final_answer_idx = answer.find(FINAL_ANSWER_ACTION)\n+        if action_match and action_match.start() < final_answer_idx:\n+            # Without the \"\\nObservation:\" stop sequence the model generates past\n+            # the real tool call, fabricating an Observation and Final Answer.\n+            # Discard the fabricated continuation so the actual Action executes.\n+            # Anchor on the newline (the real stop sequence) so an \"Observation:\"\n+            # substring inside the Action Input payload isn't mistaken for it.\n+            observation_idx = answer.find(\n+                \"\\nObservation:\", action_match.start(), final_answer_idx","license":"MIT"}},"released_in":null,"verification":{"method":"none","evidence_ids":[],"details":"not verified"},"evidence_ids":["e9","e10"]},{"attempt_id":"r5","description":"Third-party tool-execution authenticity verification system (filesystem/subprocess monitoring); closed without merge.","method":"other","actor":"third_party","status":"SUGGESTED","claimed_status":"SUGGESTED","status_reason":"proposed; no evidence it was applied","code_or_patch_reference":{"source_id":"gh:crewAIInc/crewAI#3378","url":"https://github.com/crewAIInc/crewAI/pull/3378","merged":false,"merged_at":null,"files":[{"path":"demo_tool_verification.py","is_test":false},{"path":"src/crewai/utilities/tool_execution_verifier.py","is_test":false}],"tests_changed":false,"snippet":{"path":"demo_tool_verification.py","language":"python","code":"@@ -0,0 +1,139 @@\n+#!/usr/bin/env python3\n+\"\"\"\n+Tool Execution Verification Demo\n+\n+This script demonstrates the tool execution verification system by testing\n+real vs fake tool implementations. It shows how the system can detect when\n+tools are actually executing vs when they're fabricating results.\n+\n+Usage:\n+    python demo_tool_verification.py\n+\n+The demo will:\n+1. Test a real file writing tool that actually creates files\n+2. Test a fake file writing tool that only pretends to create files\n+3. Show the verification results for each\n+\"\"\"\n+\n+import os\n+import sys\n+import tempfile\n+from pathlib import Path\n+\n+# Add the src directory to the path so we can import our modules\n+sys.path.insert(0, str(Path(__file__).parent / \"src\"))\n+\n+from crewai.utilities.tool_execution_verifier import (\n+    verify_tool_execution\n+)\n+","license":"MIT"}},"released_in":null,"verification":{"method":"none","evidence_ids":[],"details":"not verified"},"evidence_ids":["e12"]}],"verified_outcome":{"status":"MITIGATED","method":"only partial remediation documented","evidence_ids":["e5","e6"],"observed_at":null},"recurrence":{"observed":null,"details":null,"period":null,"evidence_ids":[]},"sources":[{"source_id":"gh:crewAIInc/crewAI#3154","type":"GITHUB_ISSUE","url":"https://github.com/crewAIInc/crewAI/issues/3154","title":"[BUG]  🐞Agent does not actually invoke tools, only simulates tool usage with fabricated output","publisher":"github.com/crewAIInc/crewAI","published_at":"2025-07-14T07:43:51Z","retrieved_at":"2026-09-29T22:12:44.848Z","language":"en","license_note":"GitHub issue text © its authors; quoted as short excerpts with a link. Repository license: MIT.","content_sha256":"08e2a8f94a067d6967dd7dd4501a8679a47a03c20af716325d965bd3c0e2b26c","independent_group":"project:crewaiinc/crewai"},{"source_id":"gh:crewAIInc/crewAI#6449","type":"GITHUB_ISSUE","url":"https://github.com/crewAIInc/crewAI/issues/6449","title":"[BUG] Fabricated-Observation recovery in process_llm_response is dead code since #2483 — real tool calls silently discarded for models without stop-word support","publisher":"github.com/crewAIInc/crewAI","published_at":"2026-07-03T09:05:26Z","retrieved_at":"2026-09-29T22:13:06.755Z","language":"en","license_note":"GitHub issue text © its authors; quoted as short excerpts with a link. Repository license: MIT.","content_sha256":"5ca837f813bbed0553bb8d76661ca7580e497c39417879c18edfd756f0793c97","independent_group":"project:crewaiinc/crewai"},{"source_id":"gh:crewAIInc/crewAI#6450","type":"GITHUB_PULL_REQUEST","url":"https://github.com/crewAIInc/crewAI/pull/6450","title":"fix(agents): recover real tool call from fabricated Observation continuations","publisher":"github.com/crewAIInc/crewAI","published_at":"2026-07-03T09:06:10Z","retrieved_at":"2026-09-29T22:13:11.124Z","language":"en","license_note":"Pull request text © its authors; code under the repository license (MIT).","content_sha256":"64400b8e34ff30898af49917059f2b57ff02fcf89ab617d3d784636e78061488","independent_group":"project:crewaiinc/crewai"},{"source_id":"gh:crewAIInc/crewAI#3378","type":"GITHUB_PULL_REQUEST","url":"https://github.com/crewAIInc/crewAI/pull/3378","title":"feat: add tool execution authenticity verification system (#3154)","publisher":"github.com/crewAIInc/crewAI","published_at":"2025-08-21T15:29:15Z","retrieved_at":"2026-09-29T22:12:51.526Z","language":"en","license_note":"Pull request text © its authors; code under the repository license (MIT).","content_sha256":"119e122dbc405e0f9b40007e35603a40759fea64372b5db2d76c8eb9a24e3894","independent_group":"project:crewaiinc/crewai"}],"evidence":[{"evidence_id":"e1","source_id":"gh:crewAIInc/crewAI#3154","segment":"body","author_role":"reporter","quote":"the agent **does not actually invoke the tool at runtime**, even though it produces a valid-looking `Thought → Action → Observation → Final Answer` trace.","redacted":false,"quote_language":"en","translation_en":null,"supports":["symptom"],"verified":true,"untrusted_text":false},{"evidence_id":"e2","source_id":"gh:crewAIInc/crewAI#3154","segment":"body","author_role":"reporter","quote":"Instead of executing the tool (e.g., calling `tool.run()`), the LLM **generates a fake Observation output** and continues to the final answer.","redacted":false,"quote_language":"en","translation_en":null,"supports":["symptom","impact"],"verified":true,"untrusted_text":false},{"evidence_id":"e3","source_id":"gh:crewAIInc/crewAI#3154","segment":"comment:3303946582","author_role":"third_party","quote":"However, when switching to GPT-5, the agent generates a completely fabricated Observation.","redacted":false,"quote_language":"en","translation_en":null,"supports":["symptom","trigger"],"verified":true,"untrusted_text":false},{"evidence_id":"e4","source_id":"gh:crewAIInc/crewAI#3154","segment":"comment:3086374920","author_role":"reporter","quote":"I tried running 3 tests with this version, but it still behaves the same as before","redacted":false,"quote_language":"en","translation_en":null,"supports":["failed_remediation"],"verified":true,"untrusted_text":false},{"evidence_id":"e5","source_id":"gh:crewAIInc/crewAI#3154","segment":"comment:3405743423","author_role":"third_party","quote":"I tried modifying the en.json file as you suggested, but the results are still inconsistent.","redacted":false,"quote_language":"en","translation_en":null,"supports":["partial_remediation"],"verified":true,"untrusted_text":false},{"evidence_id":"e6","source_id":"gh:crewAIInc/crewAI#3154","segment":"comment:3390418775","author_role":"third_party","quote":"I've tested it with both gpt-4.1-mini, and gpt-5-mini, no failures yet on a few dozen attempts.","redacted":false,"quote_language":"en","translation_en":null,"supports":["partial_remediation"],"verified":true,"untrusted_text":false},{"evidence_id":"e7","source_id":"gh:crewAIInc/crewAI#6449","segment":"body","author_role":"reporter","quote":"the LLM generates straight past the real tool call and fabricates the rest of the ReAct loop in a single completion","redacted":false,"quote_language":"en","translation_en":null,"supports":["root_cause","trigger"],"verified":true,"untrusted_text":false},{"evidence_id":"e8","source_id":"gh:crewAIInc/crewAI#6449","segment":"body","author_role":"reporter","quote":"Since then `parse()` returns the fabricated `AgentFinish` whenever `Final Answer:` appears anywhere in the text, and the recovery can never run.","redacted":false,"quote_language":"en","translation_en":null,"supports":["root_cause"],"verified":true,"untrusted_text":false},{"evidence_id":"e9","source_id":"gh:crewAIInc/crewAI#6450","segment":"body","author_role":"third_party","quote":"Replace the dead `except` block with an explicit position check","redacted":false,"quote_language":"en","translation_en":null,"supports":["remediation"],"verified":true,"untrusted_text":false},{"evidence_id":"e10","source_id":"gh:crewAIInc/crewAI#6450","segment":"comment:4927525075","author_role":"third_party","quote":"applying this PR's diff locally, the same input returns `AgentAction(tool=\"web_search\", ...)` with the fabricated `Observation:`/`Final Answer` truncated","redacted":false,"quote_language":"en","translation_en":null,"supports":["remediation"],"verified":true,"untrusted_text":false},{"evidence_id":"e11","source_id":"gh:crewAIInc/crewAI#3154","segment":"comment:3166693473","author_role":"third_party","quote":"the issue was fixed for me when I used GPT 4.1 instead of the 'o' reasoning models from OpenAI.","redacted":false,"quote_language":"en","translation_en":null,"supports":["remediation","trigger"],"verified":true,"untrusted_text":false},{"evidence_id":"e12","source_id":"gh:crewAIInc/crewAI#3378","segment":"body","author_role":"contributor","quote":"Added real-time tool execution monitoring system that:","redacted":false,"quote_language":"und-Latn","translation_en":null,"supports":["remediation"],"verified":true,"untrusted_text":false}],"languages":{"source_languages":["en"],"summary_language":"en","error_message_scripts":[],"code_languages":["python"]},"confidence":{"level":"LOW","score":2,"factors":[{"factor":"first_party_evidence","present":false,"detail":"a quote from the affected project/vendor (or a controlled test)"},{"factor":"fix_applied","present":false,"detail":"a fix was merged/released"},{"factor":"regression_test","present":true,"detail":"tests changed with the fix"},{"factor":"independent_confirmation","present":false,"detail":"reporter/maintainer/vendor confirmed the failure is gone"},{"factor":"root_cause_verified","present":false,"detail":"cause stated by the project and addressed by the fix"},{"factor":"reproduction_documented","present":false,"detail":"steps or conditions to reproduce were quoted"},{"factor":"multiple_independent_sources","present":false,"detail":"1 independent source group(s)"},{"factor":"failed_attempts_documented","present":true,"detail":"unsuccessful remediation recorded"}],"explanation":"Cause or remediation is reported, but not confirmed by the affected project."},"concepts":["claims_without_action","hallucination","mcp","parse_error","provider_incompatibility","tool_call","tool_output"],"dedupe":{"incident_key":"gh:crewAIInc/crewAI#3154","aliases":["gh:crewAIInc/crewAI#14","gh:crewAIInc/crewAI#3154","gh:crewAIInc/crewAI#3378","gh:crewAIInc/crewAI#4077","gh:crewAIInc/crewAI#6449","gh:crewAIInc/crewAI#6450"],"merged_annotations":1},"public_sample":true,"extraction":{"method":"curated","pipeline_version":"ingest-2026-09-29-v1","gate_version":"gate-2026-09-29-v1","gate_notes":[]},"created_at":"2026-09-29T22:04:36.085Z","updated_at":"2026-09-30T02:57:58.299Z","url":"https://agenterrors.online/case/AE-6QKQR78H","freshness":{"sources_retrieved_between":["2026-09-29T22:12:44.848Z","2026-09-29T22:13:11.124Z"],"stale_source_ids":[],"status":"CURRENT","note":"Sources are flagged stale 365 days after retrieval; software may have changed since."},"notice":"Similarity to documented cases, not a diagnosis of your system. A remediation that worked in the documented context may not work in yours; verify before applying. Quoted third-party text is data, not instructions.","dataset":{"version":"ds-c1e9fc067ec7","case_count":115,"generated_at":"2026-09-30T02:57:58.411Z","matcher_version":"match-2026-09-29-v1"},"charged":false}