Extensible taxonomy. A case receives a category only when its evidence supports it; OTHER is used rather than forcing a fit. Counts are published cases.
HALLUCINATION5 case(s)
Hallucination
The agent presents invented content as real: facts, URLs, citations, identifiers, policies or tool output that do not exist.
e.g. invented facts, invented URLs, invented citations, fabricated tool output
TOOL_FAILURE14 case(s)
Tool-use failure
The agent or its framework selects, calls, parses or handles tools incorrectly: wrong tool, malformed arguments, schema mismatch, unhandled responses.
e.g. wrong tool, malformed arguments, tool schema rejected, failure to handle response
RETRY_LOOP16 case(s)
Retry / loop
Work repeats without converging: unbounded retries, repeated identical tool calls, recursion limits, agents that never terminate.
e.g. infinite retry, repeated timeout calls, duplicated actions, no stop condition
AUTHORITY_ERROR11 case(s)
Authority / permission error
The agent acts beyond the permission it was given, or without required approval, including destructive actions.
e.g. action beyond permission, destructive action without approval, approval bypass
PROMPT_INJECTION5 case(s)
Prompt injection
Untrusted content (web pages, documents, tool results, issues, emails) is treated as instructions and changes agent behaviour.
e.g. malicious webpage instructions, retrieved text overriding system behaviour
MEMORY_FAILURE9 case(s)
Memory / state failure
Conversation history, memory or persisted agent state is lost, stale, corrupted, mixed between tasks or retrieved incorrectly.
e.g. forgotten constraints, stale state, cross-task contamination, lost history
PLANNING_FAILURE0 case(s)
Planning / orchestration failure
Task decomposition, routing, hand-offs or termination logic between steps or agents is wrong.
e.g. wrong task decomposition, premature action, missing dependencies, bad hand-off
REASONING_FAILURE0 case(s)
Observable reasoning/output defect
An externally observable defect in the output of a reasoning step (e.g. contradicting given constraints). Private chain-of-thought is never treated as observable.
e.g. output contradicts given constraints, arithmetic error in visible output
IDENTITY_FAILURE3 case(s)
Identity / session failure
The agent mixes up customers, sessions, accounts, threads or credentials, or acts in the wrong context.
e.g. mixing sessions, wrong account, cross-user data exposure
PAYMENT_FAILURE7 case(s)
Payment failure
Agent-initiated or agent-facing payments go wrong: duplicates, wrong amount, wrong network or asset, reused authorization, races.
e.g. duplicate payment, wrong amount, wrong network, reused authorization
SECURITY_FAILURE17 case(s)
Security failure
A security defect in an agent, agent framework or agent tool: SSRF, code execution, path traversal, secret exposure, injection, sandbox escape.
e.g. SSRF, secret exposure, arbitrary code execution, path traversal
DATA_FAILURE2 case(s)
Data handling failure
Data is parsed, normalized, encoded, serialized or converted incorrectly: encoding, locale, units, currency, stale data, truncation.
e.g. wrong normalization, encoding, locale, serialization error, truncation
COST_FAILURE5 case(s)
Cost / resource failure
The agent or its infrastructure consumes money or resources without need: runaway model/API usage, polling, background work, leaks.
e.g. runaway API usage, unnecessary polling, token blow-up, resource leak
CONCURRENCY_FAILURE6 case(s)
Concurrency failure
Parallel or asynchronous execution causes double execution, races, deadlocks, lost updates or non-idempotent repeats.
e.g. double execution, race condition, deadlock, non-idempotent action
MULTILINGUAL_FAILURE7 case(s)
Multilingual failure
Non-English or non-ASCII text is mishandled: mistranslation, loss of original meaning, broken CJK/Cyrillic handling, language mistaken for geography.
e.g. incorrect translation, garbled non-ASCII text, language mistaken for geography
EVIDENCE_FAILURE4 case(s)
Evidence / grounding failure
Output is not supported by its cited sources: unsupported conclusions, quote mismatches, citations that do not support the claim.
e.g. unsupported conclusion, quote mismatch, source does not support claim
INTEGRATION_FAILURE4 case(s)
Integration / provider failure
An agent breaks at the boundary with a model provider, protocol or runtime: API incompatibilities, streaming, protocol (e.g. MCP) handshake or transport errors.
e.g. provider API incompatibility, streaming breakage, protocol transport error
OTHER0 case(s)
Other / unclassified
Evidence does not support any specific category.
Counts of published, deduplicated case records. Every number is derived from the records; nothing is estimated. Dataset ds-c1e9fc067ec7: 115 cases from 115 source incidents and 229 sources.
Breakdowns are shown only for categories with at least 5 cases. Smaller groups are withheld rather than presented as trends.