AE Agent Errors

All cases

115 published, deduplicated incidents. Public summaries are free; public samples are fully open.

AE-HCGDFGC3 · OpenAI Agents SDK (Python)Invalid MCP require_approval policies silently normalize to no approval, exposing tools without human-in-the-loopAUTHORITY ERRORcause: VERIFIEDoutcome: RESOLVED UNVERIFIEDconfidence: HIGHAE-ZR5EPMXF · Vercel AI SDK (@ai-sdk/policy-opa)OPA tool-approval policy fails open: unrecognized policy decisions are treated as 'no opinion' and the protected tool executesAUTHORITY ERRORcause: LIKELYoutcome: RESOLVED VERIFIEDconfidence: LOWAE-2Q3KY4XY · Pydantic AIResumed agent run executes a requires_approval tool when the supplied approval value is NoneAUTHORITY ERRORcause: LIKELYoutcome: MITIGATEDconfidence: MEDIUMAE-8XG6KJ2H · Claude Code (@anthropic-ai/claude-code)Command-parsing error lets injected content run untrusted commands through find without the coding agent's approval promptAUTHORITY ERRORcause: VERIFIEDoutcome: RESOLVED UNVERIFIEDconfidence: MEDIUMAE-9V72XDR6 · Mastra (@mastra/core)Workspace execute_command tool sends model-chosen commands to the host shell with no approval gate by default on an un-isolated local sandboxAUTHORITY ERRORcause: LIKELYoutcome: RESOLVED UNVERIFIEDconfidence: LOWAE-JR9HK0TH · LangChainHuman-in-the-loop middleware silently drops an invalid approval config, so a tool listed for approval runs unattendedAUTHORITY ERRORcause: LIKELYoutcome: RESOLVED UNVERIFIEDconfidence: LOWAE-QXGRV1AW · Mastra (@mastra/core)Agent tool approval card in a shared chat channel can be approved by any user, not only the requesterAUTHORITY ERRORcause: LIKELYoutcome: RESOLVED UNVERIFIEDconfidence: LOWAE-R5K4NVHC · Gemini CLI (@google/gemini-cli) and run-gemini-cli GitHub ActionCoding agent in --yolo mode ignored fine-grained tool allowlists and auto-trusted workspace folders in CI, enabling code execution via prompt injectionAUTHORITY ERRORcause: VERIFIEDoutcome: RESOLVED UNVERIFIEDconfidence: MEDIUMAE-ZMZM57QY · Gemini CLIBrowser agent autonomously bypasses its allowedDomains restriction by loading a blocked site through a translation proxy on an allowed domainAUTHORITY ERRORcause: LIKELYoutcome: RESOLVED UNVERIFIEDconfidence: LOWAE-6VAFPR91 · Codex CLICoding agent on Windows force-deletes a folder with read-only files without the approval prompt that rm -rf triggers on LinuxAUTHORITY ERRORcause: LIKELYoutcome: RESOLVED UNVERIFIEDconfidence: LOWAE-95FAGZB6 · Codex CLICoding agent runs git branch deletion without an approval prompt because git branch is classified as a safe commandAUTHORITY ERRORcause: LIKELYoutcome: RESOLVED UNVERIFIEDconfidence: LOWAE-52RYEZYX · OpenAI Agents SDK (Python)Concurrent first use of an agent conversation session creates two conversations and splits the session historyCONCURRENCY FAILUREcause: VERIFIEDoutcome: RESOLVED UNVERIFIEDconfidence: HIGHAE-FPJJDB6B · LangGraph.js / LangChain.jsConcurrent invocations of a shared singleton agent leak one customer's conversation data into another threadCONCURRENCY FAILUREcause: VERIFIEDoutcome: RESOLVED UNVERIFIEDconfidence: HIGHAE-0MTWA8VG · CrewAIAgents sharing one LLM instance accumulate each other's stop words, polluting generation across agents and runsCONCURRENCY FAILUREcause: LIKELYoutcome: RESOLVED UNVERIFIEDconfidence: LOWAE-D3YS0HKK · LangGraphCompleted @task inside a nested subgraph is re-executed on every resume instead of reusing checkpointed writesCONCURRENCY FAILUREcause: LIKELYoutcome: UNKNOWNconfidence: LOWAE-M27DSV48 · OpenAI Agents SDK (Python)Concurrent first writes to a fresh SQLAlchemy agent session raise database errors and silently drop session history itemsCONCURRENCY FAILUREcause: LIKELYoutcome: RESOLVED UNVERIFIEDconfidence: LOWAE-H2M62MK2 · AutoGen AgentChatAgent issues two parallel calls to the same team tool and the second call fails because the team is already runningCONCURRENCY FAILUREcause: LIKELYoutcome: RESOLVED UNVERIFIEDconfidence: LOWAE-1441W9QY · OpenAI CodexCoding agent leaks stdio MCP server child processes across chats and shutdowns, growing process count and memory without boundCOST FAILUREcause: LIKELYoutcome: RESOLVED UNVERIFIEDconfidence: LOWAE-YVR8TSV2 · DifyStopping a streaming LLM response in a workflow/chat app leaves the provider stream running, so tokens keep being consumedCOST FAILUREcause: LIKELYoutcome: RESOLVED UNVERIFIEDconfidence: LOWAE-2BQWVVR6 · Pydantic AIStructured-output retry messages repeat the full model output in every validation error entry, inflating tokens on each retryCOST FAILUREcause: LIKELYoutcome: RESOLVED UNVERIFIEDconfidence: LOWAE-RA9EMHCP · HaystackCancelling an async agent/pipeline leaves the OpenAI streaming response open, so tokens keep being generated and chargedCOST FAILUREcause: LIKELYoutcome: RESOLVED UNVERIFIEDconfidence: LOWAE-ZJSMMDTB · Cline · public sampleCoding agent's large prompts hit provider token rate limits (429), interrupting tasks; header-aware retries added later only partly helpedCOST FAILUREcause: UNKNOWNoutcome: MITIGATEDconfidence: LOWAE-1Y0TW38H · smolagentsToolCallingAgent crashes when the model returns a tool-only turn with null content because stop-sequence trimming calls split() on NoneDATA FAILUREcause: LIKELYoutcome: RESOLVED UNVERIFIEDconfidence: LOWAE-EPEVSAZN · LlamaIndexBedrock Converse streaming stores tool-call arguments as a JSON string, breaking replay of agent history through another providerDATA FAILUREcause: LIKELYoutcome: RESOLVED UNVERIFIEDconfidence: LOWAE-H8AK9466 · Semantic Kernel (Python)Agent wrapper drops the source filename (title) from url_citation annotations, leaving citations traceable only to opaque document idsEVIDENCE FAILUREcause: VERIFIEDoutcome: RESOLVED UNVERIFIEDconfidence: HIGHAE-R0X8EHXC · Vercel AI SDK (@ai-sdk/anthropic)Web-search citations are silently stripped from assistant text when multi-turn history is replayed to the providerEVIDENCE FAILUREcause: VERIFIEDoutcome: RESOLVED UNVERIFIEDconfidence: HIGHAE-ZCA12JWN · Pydantic AI (pydantic-ai-slim)File Search citations are silently dropped on Gemini 3+ when the agent also has function toolsEVIDENCE FAILUREcause: VERIFIEDoutcome: RESOLVED UNVERIFIEDconfidence: HIGHAE-K88Y9JMS · LlamaIndex (llama-index-core)Citation query engine duplicates citation source nodes when the citation chunk size is smaller than the retrieved nodeEVIDENCE FAILUREcause: LIKELYoutcome: RESOLVED UNVERIFIEDconfidence: LOWAE-KHB97FZZ · OpenAI Agents SDK (Python)A hallucinated or unregistered tool name aborts the whole agent run with ModelBehaviorError instead of letting the model recoverHALLUCINATIONcause: UNKNOWNoutcome: RESOLVED UNVERIFIEDconfidence: LOWAE-DVYMQV5D · Pydantic AI (pydantic-ai-slim[anthropic])Model hallucinates a call to a native tool that was not enabled; replaying it on retry makes the provider reject the request with HTTP 400HALLUCINATIONcause: LIKELYoutcome: RESOLVED UNVERIFIEDconfidence: LOWAE-MNNA8GQ3 · AutoGPT Platform (AutoPilot)Assistant agent requests GitHub credentials when a Gmail block needs Google credentials (hallucinated provider in a setup tool call)HALLUCINATIONcause: LIKELYoutcome: RESOLVED UNVERIFIEDconfidence: MEDIUMAE-QKEZSMZD · Google ADK (google-adk)Agent loops calling a skill-script tool with fabricated arguments because the tool is exposed even when no skill has scriptsHALLUCINATIONcause: LIKELYoutcome: UNKNOWNconfidence: MEDIUMAE-6QKQR78H · CrewAI · public sampleReAct agent emits a tool Action but never runs the tool, fabricating the Observation and Final Answer itselfHALLUCINATIONcause: LIKELYoutcome: MITIGATEDconfidence: LOWAE-M2TJC01A · Pydantic AIShared MCP client instance sends concurrent agent runs' requests under the first caller's per-user credentialsIDENTITY FAILUREcause: LIKELYoutcome: RESOLVED UNVERIFIEDconfidence: LOWAE-63NJWB38 · AgnoPer-user agent entity memory shares one row across users with the same entity name, leaking private facts and losing dataIDENTITY FAILUREcause: LIKELYoutcome: RESOLVED UNVERIFIEDconfidence: LOWAE-ZJWTN8VB · AgnoWhatsApp interface routes messages from different phone numbers into one agent team session, leaking memories between usersIDENTITY FAILUREcause: LIKELYoutcome: RESOLVED UNVERIFIEDconfidence: LOWAE-3FRMS134 · LangGraphAfter upgrading LangGraph, interrupt() inside an agent tool is not surfaced when streaming with stream_mode="values"INTEGRATION FAILUREcause: VERIFIEDoutcome: RESOLVED UNVERIFIEDconfidence: HIGHAE-4J03X6FG · AgnoAgent session replay reconstructs Claude assistant turns lossily, so Anthropic rejects modified thinking blocks and resumed runs stay stuckINTEGRATION FAILUREcause: VERIFIEDoutcome: RESOLVED UNVERIFIEDconfidence: HIGHAE-HV7EDV39 · MCP filesystem server · public sampleFilesystem MCP server returns structuredContent that violates its own outputSchema, so the coding agent's MCP client rejects directory_tree resultsINTEGRATION FAILUREcause: VERIFIEDoutcome: RESOLVED UNVERIFIEDconfidence: HIGHAE-NZFG9GC5 · Semantic Kernel (Python)Gemini 3 function calling with thinking enabled fails with HTTP 400 because the connector drops thought_signature on replayed function callsINTEGRATION FAILUREcause: LIKELYoutcome: RESOLVED UNVERIFIEDconfidence: LOWAE-GWWP1DFD · Model Context Protocol reference servers (memory)Memory MCP server corrupts or loses its knowledge-graph file when an agent issues several memory tool calls in one turnMEMORY FAILUREcause: VERIFIEDoutcome: RESOLVED UNVERIFIEDconfidence: HIGHAE-MY5FEYWV · Vercel AI SDK (@ai-sdk/harness-opencode)Resumed coding-agent harness session returns the previous turn's answer verbatim as a successful reply to a new promptMEMORY FAILUREcause: LIKELYoutcome: RESOLVED UNVERIFIEDconfidence: MEDIUMAE-BTYX7BWF · CAMEL · public sampleAgent memory orders a tool result before its tool call when both share a timestamp, so the next model request is rejectedMEMORY FAILUREcause: LIKELYoutcome: RESOLVED UNVERIFIEDconfidence: MEDIUMAE-JKGSV5K4 · MemGPT (now Letta)Interrupted or failed agent creation leaves a stateless agent directory that can no longer be loadedMEMORY FAILUREcause: LIKELYoutcome: RESOLVED UNVERIFIEDconfidence: MEDIUMAE-R2SQ9HX4 · Semantic Kernel (Python)Chat history truncation removes the assistant tool-call message and leaves orphaned tool results, so the agent thread is rejected by the APIMEMORY FAILUREcause: LIKELYoutcome: RESOLVED UNVERIFIEDconfidence: LOWAE-T35FP6NY · LangChainSummarization middleware deletes agent conversation history and stores the provider error text as the summary when the summary call failsMEMORY FAILUREcause: LIKELYoutcome: RESOLVED UNVERIFIEDconfidence: LOWAE-C6TC6MGJ · smolagentsAgent memory in the Gradio UI is shared by all users, so one user's long chat exhausts the context window for everyoneMEMORY FAILUREcause: LIKELYoutcome: RESOLVED UNVERIFIEDconfidence: LOWAE-0W7Q01W1 · LlamaIndex (llama-index-core)Agent memory stores chat messages out of order when messages are added in rapid successionMEMORY FAILUREcause: LIKELYoutcome: RESOLVED UNVERIFIEDconfidence: LOWAE-9PXHKDBW · LangGraph.js / LangChain.jsHuman-approval graph breaks on the next invocation after a tool call is rejected without a matching tool messageMEMORY FAILUREcause: LIKELYoutcome: UNKNOWNconfidence: LOWAE-64GJNA0D · shopping agent (controlled sample) · controlled testAgent infers the quote currency from the message language instead of the customer's countryMULTILINGUAL FAILUREcause: VERIFIEDoutcome: RESOLVED VERIFIEDconfidence: HIGHAE-VH5D57DM · tool-calling agent (controlled sample) · controlled testByte-based truncation of agent output splits a Japanese character into a replacement characterMULTILINGUAL FAILUREcause: VERIFIEDoutcome: RESOLVED VERIFIEDconfidence: HIGHAE-3ZDD3QWC · LiteLLMGemini tool-call arguments returned through an LLM gateway escape Japanese text as \uXXXX sequencesMULTILINGUAL FAILUREcause: LIKELYoutcome: RESOLVED VERIFIEDconfidence: LOWAE-YPSPX8EQ · @modelcontextprotocol/server-filesystemMCP filesystem server head/tail reads corrupt CJK characters that straddle a 1024-byte chunk boundaryMULTILINGUAL FAILUREcause: LIKELYoutcome: RESOLVED UNVERIFIEDconfidence: MEDIUMAE-AZ76K1SH · Google ADK (adk-python)Agent answers an English-speaking user in Russian after large tool responses in a long-running sessionMULTILINGUAL FAILUREcause: LIKELYoutcome: RESOLVED UNVERIFIEDconfidence: MEDIUMAE-CY956VR0 · Codex CLICoding agent sends raw non-ASCII workspace path in an HTTP header, so requests and every reconnect fail for Japanese/Chinese pathsMULTILINGUAL FAILUREcause: LIKELYoutcome: RESOLVED UNVERIFIEDconfidence: LOWAE-BQWKM2WQ · Gemini CLICoding agent's shell tool returns garbled Japanese (mojibake) in command output on WindowsMULTILINGUAL FAILUREcause: LIKELYoutcome: UNKNOWNconfidence: LOWAE-ZEKX4YYA · purchasing agent (controlled sample) · controlled testAgent retries a purchase after a lost response and is charged twicePAYMENT FAILUREcause: VERIFIEDoutcome: RESOLVED VERIFIEDconfidence: HIGHAE-HHZNJ5QC · x402 (@x402/core)x402 settlement override can resolve above the payer's authorized maximum and is forwarded to settle without errorPAYMENT FAILUREcause: LIKELYoutcome: UNKNOWNconfidence: LOWAE-YY2AK4A7 · x402 (@x402/fetch client, facilitator)x402 payment settles on-chain after the facilitator times out, so the payer is debited but the paid request is rejectedPAYMENT FAILUREcause: LIKELYoutcome: RESOLVED UNVERIFIEDconfidence: LOWAE-HMETQYGG · Coinbase AgentKit (coinbase-agentkit)Agent vault-withdraw action converts amounts with a hardcoded 18 decimals, off by 10^12 for 6-decimal tokens such as USDCPAYMENT FAILUREcause: LIKELYoutcome: UNKNOWNconfidence: LOWAE-N4TPJBDY · Coinbase AgentKitAgent ERC-20 transfer action does not scale decimal amounts to token units, so transferring 0.01 USDC reverts or failsPAYMENT FAILUREcause: LIKELYoutcome: UNKNOWNconfidence: LOWAE-TA1EJAB1 · awal (agent wallet CLI)Agent wallet CLI x402 pay authorizes payment but resubmits on GET, so a POST-only settlement endpoint rejects itPAYMENT FAILUREcause: LIKELYoutcome: UNKNOWNconfidence: LOWAE-EE1QN5D2 · Stripe agent toolkit / MCP (stripe/ai)Agent refund tool call fails with 'Invalid integer: 1000.0' because the MCP tool schema types the amount as a numberPAYMENT FAILUREcause: LIKELYoutcome: UNKNOWNconfidence: LOWAE-97N3CYDM · LlamaIndexPrompt injection in a pandas query engine makes the LLM emit Python that is run with exec, giving code execution (CVE-2023-39662)PROMPT INJECTIONcause: LIKELYoutcome: RESOLVED UNVERIFIEDconfidence: MEDIUMAE-CFMQS0WN · Claude Code (@anthropic-ai/claude-code)Coding agent's overly broad safe-command allowlist lets injected instructions read a file and send it over the network without confirmationPROMPT INJECTIONcause: VERIFIEDoutcome: RESOLVED UNVERIFIEDconfidence: MEDIUMAE-H4RRJ172 · Claude Code (@anthropic-ai/claude-code)Coding agent's web-fetch tool auto-approves any path on a pre-approved domain, letting injected content exfiltrate data through download countersPROMPT INJECTIONcause: VERIFIEDoutcome: RESOLVED UNVERIFIEDconfidence: MEDIUMAE-4FBXM5ZK · GitHub MCP server with Claude DesktopMalicious public GitHub issue hijacks an agent using the GitHub MCP server into leaking private repository data via a public pull requestPROMPT INJECTIONcause: LIKELYoutcome: UNKNOWNconfidence: LOWAE-PYMEWZQX · GitLab DuoHidden prompts in project content make a code assistant leak private source code through HTML injected into its streamed answerPROMPT INJECTIONcause: LIKELYoutcome: UNKNOWNconfidence: LOWAE-C9JF11Z2 · tool-calling agent (controlled sample) · public sample · controlled testTool wrapper retries a timing-out tool call without any limitRETRY LOOPcause: VERIFIEDoutcome: RESOLVED VERIFIEDconfidence: HIGHAE-JQDA71M6 · OpenAI Agents SDK (Python)Nested agent-as-tool with approval-gated tool pauses again on every resume after RunState JSON round-trip, causing an endless approve/resume cycleRETRY LOOPcause: VERIFIEDoutcome: RESOLVED UNVERIFIEDconfidence: HIGHAE-72B9RH99 · LangChain.jsAll HTTP 429 responses treated as retryable, so quota-exhaustion errors are retried repeatedly and fallbacks never triggerRETRY LOOPcause: VERIFIEDoutcome: RESOLVED UNVERIFIEDconfidence: HIGHAE-7TY99VF7 · Vercel AI SDKChat client enters an infinite client/server request loop when a tool call errors and auto-send is enabledRETRY LOOPcause: VERIFIEDoutcome: RESOLVED UNVERIFIEDconfidence: HIGHAE-A5CVCAQJ · Semantic Kernel (Python)OpenAI Assistants agent polls a run stuck in 'incomplete' status indefinitely because the status was not treated as terminal and no timeout existedRETRY LOOPcause: VERIFIEDoutcome: RESOLVED UNVERIFIEDconfidence: HIGHAE-01V5S060 · LiveKit AgentsVoice-agent LLM plugin retries at two layers (vendor SDK plus framework), so attempt timeouts are exceeded and failover is delayedRETRY LOOPcause: LIKELYoutcome: RESOLVED UNVERIFIEDconfidence: LOWAE-2RBJEHVQ · Pydantic AIPer-tool retry budget resets when the model interleaves another tool call, so an always-failing tool is retried past max_retriesRETRY LOOPcause: LIKELYoutcome: RESOLVED UNVERIFIEDconfidence: LOWAE-CQNQB0VG · Vercel AI SDKChat client resends messages forever after provider-executed web search/fetch tool calls when auto-send predicate stays true on empty repliesRETRY LOOPcause: LIKELYoutcome: RESOLVED UNVERIFIEDconfidence: LOWAE-FN92K6QP · HaystackAgent re-calls a reasoning model until max_agent_steps when every reply is empty because reasoning exhausts max_output_tokensRETRY LOOPcause: LIKELYoutcome: RESOLVED UNVERIFIEDconfidence: LOWAE-5Z8KSAZF · LlamaIndexReAct agent overshoots max_iterations without stopping because the limit check uses equality while several reasoning steps can be added per iterationRETRY LOOPcause: LIKELYoutcome: RESOLVED UNVERIFIEDconfidence: LOWAE-A42PNS2K · Google ADK (Python)LLM agent with tools and include_contents="none" re-calls its tool in an endless loop because tool results never reach the next model callRETRY LOOPcause: LIKELYoutcome: RESOLVED UNVERIFIEDconfidence: LOWAE-05NE9DC4 · Gemini CLICoding-agent CLI's background prompt completion retries a failing model request in an endless loop (e.g. on quota exceeded)RETRY LOOPcause: LIKELYoutcome: RESOLVED UNVERIFIEDconfidence: LOWAE-AFGNRN1C · LangGraph / LangChain · public sampleAgent with structured response format loops on a failing SQL tool until the recursion limitRETRY LOOPcause: UNKNOWNoutcome: UNKNOWNconfidence: LOWAE-07ZGFCMV · DifyAgent node in a chatflow repeatedly calls the same MCP tool with an injected "continue" query instead of ending the turnRETRY LOOPcause: LIKELYoutcome: UNKNOWNconfidence: LOWAE-HJEA45BZ · LangGraphGraph agent keeps looping after its router returns END because an unconditional edge also leaves the same nodeRETRY LOOPcause: LIKELYoutcome: UNKNOWNconfidence: LOWAE-RY51XRQG · LangChainPandas dataframe agent loops on a nonexistent tool name until the iteration limit instead of using its Python toolRETRY LOOPcause: UNKNOWNoutcome: UNKNOWNconfidence: LOWAE-Q8D6VCC7 · browsing agent (controlled sample) · controlled testURL-fetching tool validates only the first URL and follows a redirect to the cloud metadata addressSECURITY FAILUREcause: VERIFIEDoutcome: RESOLVED VERIFIEDconfidence: HIGHAE-GDZEYNCQ · LlamaIndexRestricted safe_eval of LLM-generated code is bypassed via prompt injection, executing arbitrary code (bypass of an earlier fix)SECURITY FAILUREcause: LIKELYoutcome: RESOLVED UNVERIFIEDconfidence: MEDIUMAE-RQNNT84C · Langflow (uses LangChain create_csv_agent)CSV agent node hardcodes allow_dangerous_code, exposing a Python REPL tool that runs prompt-injected code on the serverSECURITY FAILUREcause: VERIFIEDoutcome: RESOLVED UNVERIFIEDconfidence: HIGHAE-6S0G2VMH · Browser UseBrowser agent's allowed_domains check is bypassed by placing an allowed domain in the URL's credentials partSECURITY FAILUREcause: VERIFIEDoutcome: RESOLVED UNVERIFIEDconfidence: HIGHAE-7D9R8W4R · Pydantic AIAgent framework downloads file URLs from untrusted message history without address checks, reaching internal services and cloud metadataSECURITY FAILUREcause: VERIFIEDoutcome: RESOLVED UNVERIFIEDconfidence: HIGHAE-AN41MM0F · mcp-remoteMCP client proxy executes OS commands from a crafted authorization_endpoint URL returned by an untrusted MCP serverSECURITY FAILUREcause: LIKELYoutcome: RESOLVED UNVERIFIEDconfidence: MEDIUMAE-BFPF6JD0 · LangChainUnescaped 'lc' keys in serialized data let prompt-injected LLM response fields load environment secrets on deserializationSECURITY FAILUREcause: VERIFIEDoutcome: RESOLVED UNVERIFIEDconfidence: HIGHAE-BNQTVDWF · smolagentsCode agent's local Python executor sandbox is escaped through whitelisted modules and functions, giving arbitrary code executionSECURITY FAILUREcause: LIKELYoutcome: RESOLVED UNVERIFIEDconfidence: MEDIUMAE-GR7ZDC99 · OpenAI Codex CLICoding agent treats a model-generated cwd as the sandbox's writable root, allowing writes and commands outside the workspaceSECURITY FAILUREcause: VERIFIEDoutcome: RESOLVED UNVERIFIEDconfidence: HIGHAE-YKAMK5PX · Model Context Protocol reference servers (Filesystem) · public sampleFilesystem MCP server grants access to files in directories whose path only shares a prefix with an allowed directorySECURITY FAILUREcause: VERIFIEDoutcome: RESOLVED UNVERIFIEDconfidence: HIGHAE-9C54Z2RQ · Claude CodeCoding agent's unsandboxed writer follows a symlink created by a sandboxed command, writing files outside the workspaceSECURITY FAILUREcause: VERIFIEDoutcome: RESOLVED UNVERIFIEDconfidence: MEDIUMAE-BW0X6WZ1 · LangChainAgent file-search middleware and config loaders let glob patterns, symlinks and prefix checks reach files outside the configured rootSECURITY FAILUREcause: VERIFIEDoutcome: RESOLVED UNVERIFIEDconfidence: MEDIUMAE-E1MM1WST · n8nAI agent MCP connector ignores a credential's allowed-domain restriction, sending the shared secret to an arbitrary URLSECURITY FAILUREcause: VERIFIEDoutcome: RESOLVED UNVERIFIEDconfidence: MEDIUMAE-GNRV7889 · LangSmith SDKPulling a public Hub prompt deserializes attacker-controlled manifests, allowing model base_url redirection and secret disclosureSECURITY FAILUREcause: VERIFIEDoutcome: RESOLVED UNVERIFIEDconfidence: MEDIUMAE-K9VYGJHV · n8n (@n8n/computer-use)Computer-use agent's shell tool runs unsandboxed on Linux and Windows because sandbox restrictions were applied only on macOSSECURITY FAILUREcause: VERIFIEDoutcome: RESOLVED UNVERIFIEDconfidence: MEDIUMAE-STJ4JWMS · LiteLLMMCP server preview endpoints spawn a request-supplied stdio command on the proxy host for any authenticated keySECURITY FAILUREcause: VERIFIEDoutcome: RESOLVED UNVERIFIEDconfidence: MEDIUMAE-G3HH1AH0 · MetaGPTLLM-callable terminal tool executes arbitrary shell commands because its blocklist filters only two stringsSECURITY FAILUREcause: LIKELYoutcome: UNKNOWNconfidence: LOWAE-49973DV6 · LangChainAgent's invalid-tool-call repair leaves an orphaned tool_result, so Anthropic-backed threads fail with a permanent 400TOOL FAILUREcause: VERIFIEDoutcome: RESOLVED UNVERIFIEDconfidence: HIGHAE-1D774RZA · MastraAgent file-read tool truncates output inside an emoji, emitting a lone UTF-16 surrogate that a strict model gateway rejectsTOOL FAILUREcause: LIKELYoutcome: RESOLVED UNVERIFIEDconfidence: LOWAE-6DS0JGFN · Semantic Kernel (.NET; Python tracked separately)Auto function calling fails when the model hallucinates a function name with a disallowed separator and the name is echoed back to the APITOOL FAILUREcause: LIKELYoutcome: RESOLVED UNVERIFIEDconfidence: MEDIUMAE-QVPSQBJA · MCP TypeScript SDKMCP server rejects agent tool calls that omit arguments for tools whose parameters are all optionalTOOL FAILUREcause: LIKELYoutcome: RESOLVED UNVERIFIEDconfidence: LOWAE-WKCR0NF0 · crewAIStreaming tool calls through the Bedrock provider reach the agent's tool with empty argumentsTOOL FAILUREcause: LIKELYoutcome: RESOLVED UNVERIFIEDconfidence: LOWAE-YTT3MEGG · crewAIAzure streaming merges parallel tool calls into one corrupted call with mixed id/name and concatenated invalid JSON argumentsTOOL FAILUREcause: LIKELYoutcome: RESOLVED UNVERIFIEDconfidence: LOWAE-3KV5B4YT · AutoGenAssistantAgent fails with 422 after multiple tool calls because replayed assistant tool-call messages omit the content fieldTOOL FAILUREcause: LIKELYoutcome: RESOLVED UNVERIFIEDconfidence: LOWAE-9W7S1V4A · OpenAI Agents SDK (Python)Streamed tool-call turns from Chat Completions providers replay as an invalid message sequence, so every later agent turn fails with 400TOOL FAILUREcause: LIKELYoutcome: RESOLVED UNVERIFIEDconfidence: LOWAE-FBZ9M6NV · LettaStateful agent shows inner monologue but no user reply because the model response lacks the send_message tool call; fix did not resolve it for usersTOOL FAILUREcause: UNKNOWNoutcome: UNRESOLVEDconfidence: LOWAE-H96FWV84 · browser-useBrowser agent stores its past tool calls as plain JSON text in history, so local models imitate that format and emit malformed tool callsTOOL FAILUREcause: LIKELYoutcome: RESOLVED UNVERIFIEDconfidence: MEDIUMAE-P0BMHWX9 · Pydantic AICohere model adapter silently drops zero-argument tool calls, so the agent never executes the toolTOOL FAILUREcause: LIKELYoutcome: RESOLVED UNVERIFIEDconfidence: LOWAE-1KT9WZ9P · LangChainStructured-chat agent with a hand-written ReAct prompt raises an output parsing error on a Wikipedia tool stepTOOL FAILUREcause: UNKNOWNoutcome: UNKNOWNconfidence: LOWAE-1TPK4F13 · MetaGPTMulti-agent PRD step with DeepSeek-R1 retries until RetryError because structured output misses required fieldsTOOL FAILUREcause: UNKNOWNoutcome: UNKNOWNconfidence: LOWAE-747H0GWD · LangChainSQL database agent with a chat model fails mid-run with OutputParserException on a malformed Action lineTOOL FAILUREcause: UNKNOWNoutcome: UNKNOWNconfidence: LOW