AE Agent Errors

Error taxonomy

Extensible taxonomy. A case receives a category only when its evidence supports it; OTHER is used rather than forcing a fit. Counts are published cases.

HALLUCINATION5 case(s)

Hallucination

The agent presents invented content as real: facts, URLs, citations, identifiers, policies or tool output that do not exist.

e.g. invented facts, invented URLs, invented citations, fabricated tool output

TOOL_FAILURE14 case(s)

Tool-use failure

The agent or its framework selects, calls, parses or handles tools incorrectly: wrong tool, malformed arguments, schema mismatch, unhandled responses.

e.g. wrong tool, malformed arguments, tool schema rejected, failure to handle response

RETRY_LOOP16 case(s)

Retry / loop

Work repeats without converging: unbounded retries, repeated identical tool calls, recursion limits, agents that never terminate.

e.g. infinite retry, repeated timeout calls, duplicated actions, no stop condition

AUTHORITY_ERROR11 case(s)

Authority / permission error

The agent acts beyond the permission it was given, or without required approval, including destructive actions.

e.g. action beyond permission, destructive action without approval, approval bypass

PROMPT_INJECTION5 case(s)

Prompt injection

Untrusted content (web pages, documents, tool results, issues, emails) is treated as instructions and changes agent behaviour.

e.g. malicious webpage instructions, retrieved text overriding system behaviour

MEMORY_FAILURE9 case(s)

Memory / state failure

Conversation history, memory or persisted agent state is lost, stale, corrupted, mixed between tasks or retrieved incorrectly.

e.g. forgotten constraints, stale state, cross-task contamination, lost history

PLANNING_FAILURE0 case(s)

Planning / orchestration failure

Task decomposition, routing, hand-offs or termination logic between steps or agents is wrong.

e.g. wrong task decomposition, premature action, missing dependencies, bad hand-off

REASONING_FAILURE0 case(s)

Observable reasoning/output defect

An externally observable defect in the output of a reasoning step (e.g. contradicting given constraints). Private chain-of-thought is never treated as observable.

e.g. output contradicts given constraints, arithmetic error in visible output

IDENTITY_FAILURE3 case(s)

Identity / session failure

The agent mixes up customers, sessions, accounts, threads or credentials, or acts in the wrong context.

e.g. mixing sessions, wrong account, cross-user data exposure

PAYMENT_FAILURE7 case(s)

Payment failure

Agent-initiated or agent-facing payments go wrong: duplicates, wrong amount, wrong network or asset, reused authorization, races.

e.g. duplicate payment, wrong amount, wrong network, reused authorization

SECURITY_FAILURE17 case(s)

Security failure

A security defect in an agent, agent framework or agent tool: SSRF, code execution, path traversal, secret exposure, injection, sandbox escape.

e.g. SSRF, secret exposure, arbitrary code execution, path traversal

DATA_FAILURE2 case(s)

Data handling failure

Data is parsed, normalized, encoded, serialized or converted incorrectly: encoding, locale, units, currency, stale data, truncation.

e.g. wrong normalization, encoding, locale, serialization error, truncation

COST_FAILURE5 case(s)

Cost / resource failure

The agent or its infrastructure consumes money or resources without need: runaway model/API usage, polling, background work, leaks.

e.g. runaway API usage, unnecessary polling, token blow-up, resource leak

CONCURRENCY_FAILURE6 case(s)

Concurrency failure

Parallel or asynchronous execution causes double execution, races, deadlocks, lost updates or non-idempotent repeats.

e.g. double execution, race condition, deadlock, non-idempotent action

MULTILINGUAL_FAILURE7 case(s)

Multilingual failure

Non-English or non-ASCII text is mishandled: mistranslation, loss of original meaning, broken CJK/Cyrillic handling, language mistaken for geography.

e.g. incorrect translation, garbled non-ASCII text, language mistaken for geography

EVIDENCE_FAILURE4 case(s)

Evidence / grounding failure

Output is not supported by its cited sources: unsupported conclusions, quote mismatches, citations that do not support the claim.

e.g. unsupported conclusion, quote mismatch, source does not support claim

INTEGRATION_FAILURE4 case(s)

Integration / provider failure

An agent breaks at the boundary with a model provider, protocol or runtime: API incompatibilities, streaming, protocol (e.g. MCP) handshake or transport errors.

e.g. provider API incompatibility, streaming breakage, protocol transport error

OTHER0 case(s)

Other / unclassified

Evidence does not support any specific category.

Statistics

Counts of published, deduplicated case records. Every number is derived from the records; nothing is estimated. Dataset ds-c1e9fc067ec7: 115 cases from 115 source incidents and 229 sources.

ConfidenceCount
LOW65
HIGH26
MEDIUM24
Root causeCount
LIKELY71
VERIFIED36
UNKNOWN8
OutcomeCount
RESOLVED_UNVERIFIED85
UNKNOWN19
RESOLVED_VERIFIED7
MITIGATED3
UNRESOLVED1
Remediation attemptsCount
SUGGESTED72
TESTED62
APPLIED26
FAILED9
VERIFIED_SUCCESS7
PARTIAL_SUCCESS6
UNKNOWN1
Source type (cases)Count
GITHUB_ISSUE83
GITHUB_PULL_REQUEST81
SECURITY_ADVISORY20
GITHUB_COMMIT8
CONTROLLED_TEST5
DEVELOPER_DISCUSSION5
RESEARCH_REPORT2
Source language (cases)Count
en114
und-Latn7
ja3
zh3
Case typeCount
REAL_WORLD110
CONTROLLED_TEST_CASE5

Failure patterns

Breakdowns are shown only for categories with at least 5 cases. Smaller groups are withheld rather than presented as trends.

SECURITY_FAILURE — 17 observed cases

Root causeCount
VERIFIED13
LIKELY4
Remediation methodCount
sandboxing4
path_validation3
ssrf_protection3
input_validation2
permission_check2
url_validation2
workaround2
config_change1
output_sanitization1
Remediation statusCount
TESTED9
APPLIED6
SUGGESTED3
PARTIAL_SUCCESS1
VERIFIED_SUCCESS1

RETRY_LOOP — 16 observed cases

Root causeCount
LIKELY9
VERIFIED5
UNKNOWN2
Remediation methodCount
termination_condition7
bounded_retry5
state_persistence3
workaround3
prompt_change2
config_change1
exponential_backoff1
recursion_limit1
tool_call_limit1
Remediation statusCount
SUGGESTED10
TESTED10
FAILED2
APPLIED1
VERIFIED_SUCCESS1

TOOL_FAILURE — 14 observed cases

Root causeCount
LIKELY9
UNKNOWN4
VERIFIED1
Remediation methodCount
parser_fix5
schema_fix5
config_change4
workaround3
input_validation2
prompt_change2
dependency_upgrade1
error_handling1
output_sanitization1
Remediation statusCount
SUGGESTED13
TESTED8
APPLIED2
FAILED1

AUTHORITY_ERROR — 11 observed cases

Root causeCount
LIKELY8
VERIFIED3
Remediation methodCount
permission_check7
input_validation3
human_approval_gate1
prompt_change1
url_validation1
workaround1
Remediation statusCount
TESTED6
APPLIED4
SUGGESTED2
PARTIAL_SUCCESS1
VERIFIED_SUCCESS1

MEMORY_FAILURE — 9 observed cases

Root causeCount
LIKELY8
VERIFIED1
Remediation methodCount
workaround5
error_handling4
state_persistence3
concurrency_lock2
cache_invalidation1
config_change1
schema_fix1
stale_state_fence1
Remediation statusCount
SUGGESTED8
APPLIED4
TESTED4
FAILED2

MULTILINGUAL_FAILURE — 7 observed cases

Root causeCount
LIKELY5
VERIFIED2
Remediation methodCount
encoding_fix7
prompt_change3
workaround3
input_validation1
Remediation statusCount
SUGGESTED6
VERIFIED_SUCCESS3
APPLIED2
FAILED1
PARTIAL_SUCCESS1
TESTED1

PAYMENT_FAILURE — 7 observed cases

Root causeCount
LIKELY6
VERIFIED1
Remediation methodCount
input_validation5
idempotency_key2
workaround2
error_handling1
schema_fix1
Remediation statusCount
SUGGESTED8
TESTED2
VERIFIED_SUCCESS1

CONCURRENCY_FAILURE — 6 observed cases

Root causeCount
LIKELY4
VERIFIED2
Remediation methodCount
session_isolation5
concurrency_lock2
config_change2
other1
state_persistence1
workaround1
Remediation statusCount
SUGGESTED5
TESTED5
APPLIED1
UNKNOWN1

COST_FAILURE — 5 observed cases

Root causeCount
LIKELY4
UNKNOWN1
Remediation methodCount
error_handling2
workaround2
config_change1
exponential_backoff1
other1
output_sanitization1
Remediation statusCount
TESTED4
PARTIAL_SUCCESS2
SUGGESTED2

HALLUCINATION — 5 observed cases

Root causeCount
LIKELY4
UNKNOWN1
Remediation methodCount
config_change2
output_sanitization2
parser_fix2
prompt_change2
workaround2
error_handling1
other1
tool_call_limit1
Remediation statusCount
SUGGESTED7
FAILED2
TESTED2
APPLIED1
PARTIAL_SUCCESS1

PROMPT_INJECTION — 5 observed cases

Root causeCount
LIKELY3
VERIFIED2
Remediation methodCount
permission_check3
sandboxing2
output_sanitization1
Remediation statusCount
APPLIED2
SUGGESTED2
FAILED1
TESTED1