Coding agent's large prompts hit provider token rate limits (429), interrupting tasks; header-aware retries added later only partly helped
Users of a coding agent with the Anthropic API frequently hit 429 per-minute token rate limits even on small edits, while the agent's displayed token counts were far lower than provider usage; one user measured about 35,000 tokens of prompt before any code was added. Tasks were interrupted and the manual retry restarted the task. A merged change added automatic 429 retries with header-based timing and exponential backoff, but users still reported 429 errors in later versions.
- Framework / agent
- Cline · IDE coding agent
- Remediation attempts
- PARTIAL SUCCESSPARTIAL SUCCESS
- Recurrence
- not documented
- Source languages
- en
- Updated
- 2026-09-30
Sources
- GITHUB ISSUE 429 {"type":"error","error":{"type":"rate_limit_error","message":"Number of request tokens has exceeded your per-minute rate limit in cline when it has not been used for days. — github.com/cline/cline, retrieved 2026-09-29
- GITHUB PULL REQUEST feat: add retry decorator with rate limit handling — github.com/cline/cline, retrieved 2026-09-29
Public sample: the full evidence record is free.
Symptoms
- Simple tasks immediately fail with a 429 per-minute request-token rate-limit error REPORTED CLAIM
- The agent's prompt is about 35,000 tokens before any code is added, including full paths of all project files REPORTED CLAIM
- The retry button restarts the task instead of continuing, repeatedly consuming money REPORTED CLAIM
Error messages (verbatim)
Number of request tokens has exceeded your per-minute rate limit
Context and trigger
- Agent type
- IDE coding agent
- Component
- provider API request handling / retry; prompt construction
- Framework
- Cline
- Model
- Claude 3.5 Sonnet (reported)
- Task
- code editing tasks
- Using a provider account with low per-minute token limits (e.g. lower Anthropic tiers) while the agent sends large prompts on every request REPORTED CLAIM
Open questions
- Why the displayed token count differed from provider-reported usage was not explained in the sources.
- Whether the post-fix 429 reports come from exhausted retries or other causes was not investigated.
Root cause UNKNOWN
Not established by the sources.
no verified evidence about the cause
Remediation attempts (2)
User workaround: switch from the Anthropic API to Claude via GCP Vertex API; made the agent usable at higher cost, but 429 errors still occurred occasionally.
Status basis: documented as only partly fixing the failure. Verification: not verified.
I've switched to using GCP Vertex API (still Claude model) from Anthropic. It costs my company more, but it makes Cline usable (it isn't usable with Anthropic Tier 1 service, IMHO). Although less common, I still receive the following errorAdd a retry decorator to all providers that retries only 429 errors, using retry-after / rate-limit reset headers when present and exponential backoff otherwise (max 3 retries). Merged with unit tests; later users still reported frequent 429 errors.
Status basis: documented as only partly fixing the failure. Verification: not verified.
Change: https://github.com/cline/cline/pull/1605 · merged · tests changed
Excerpt of the change (.changeset/modern-knives-tan.md, Apache-2.0):
@@ -0,0 +1,5 @@ +--- +"claude-dev": patch +--- + +Add automatic retry for rate limited requests
Added a new `@[user]` decorator that automatically handles 429 (rate limit) errors
I don't think this is completely fixed. I'm using version 3.4.5 and I'm getting this error.Using clinebot v3.8.5 with Bedrock/Claude v3.7 and Cline almost unusable due to "429 Too many tokens" errors. I am not clear why this issue is closed.Outcome outcome: MITIGATED
only partial remediation documented
Recurrence
Not documented in the sources (absence of reports is not evidence of absence).
Confidence LOW
Cause or remediation is reported, but not confirmed by the affected project.
| Factor | Present | Meaning |
|---|---|---|
| first_party_evidence | no | a quote from the affected project/vendor (or a controlled test) |
| fix_applied | no | a fix was merged/released |
| regression_test | yes | tests changed with the fix |
| independent_confirmation | no | reporter/maintainer/vendor confirmed the failure is gone |
| root_cause_verified | no | cause stated by the project and addressed by the fix |
| reproduction_documented | no | steps or conditions to reproduce were quoted |
| multiple_independent_sources | no | 1 independent source group(s) |
| failed_attempts_documented | no | unsuccessful remediation recorded |
All evidence (7 verified quotes)
429 {"type":"error","error":{"type":"rate_limit_error","message":"Number of request tokens has exceeded your per-minute rate limit
The prompt cline is sending without any code added is over 135,000 chars or 35,000 tokens before it even gets to working with code. It includes the full path of all files in my project.The most stupid aspetto for this story is that the retry button doesn’t continue the task, but it starts again. So it generates ad infinite loop and money consuming.I've switched to using GCP Vertex API (still Claude model) from Anthropic. It costs my company more, but it makes Cline usable (it isn't usable with Anthropic Tier 1 service, IMHO). Although less common, I still receive the following errorAdded a new `@[user]` decorator that automatically handles 429 (rate limit) errors
I don't think this is completely fixed. I'm using version 3.4.5 and I'm getting this error.Using clinebot v3.8.5 with Bedrock/Claude v3.7 and Cline almost unusable due to "429 Too many tokens" errors. I am not clear why this issue is closed.Categories: COST FAILURE RETRY LOOP · extraction curated, gate-2026-09-29-v1
Similarity to your system is not implied. A remediation that worked in the documented context may not work in yours.