Anthropic SDK max_retries Stacked on My Retries: 36 Calls per Job
My application logs said the report worker called Claude 1,302 times in 22 minutes. An httpx hook I added the next morning said it was 3,904 HTTP requests. Both were right, which was the problem. That gap came from the
My application logs said the report worker called Claude 1,302 times in 22 minutes. An httpx hook I added the next morning said it was 3,904 HTTP requests. Both were right, which was the problem.
That gap came from the Anthropic SDK max_retries default sitting underneath two other retry layers I had written myself. Each layer looked reasonable on its own. Together they multiplied: 3 Γ 4 Γ 3 = 36 attempts for one failing call. During a short stretch of 529 overloaded_error responses, that multiplication turned a blip into a retry storm. When the API recovered, the storm hit my own rate limit and took down live traffic.
TL;DR
- The Anthropic Python SDK retries failed requests 2 times by default (3 attempts total) on connection errors, 408, 409, 429 and 5xx, including 529.
- If you wrap
messages.create()in tenacity, and your job queue also retries, the attempts multiply, not add. Mine was 3 Γ 4 Γ 3 = 36. - My 22-minute overload window produced 9.3Γ normal request volume from background jobs, and the recovery burst caused 429s on 18% of live turns.
- Fix: write down the product of all retry layers, let the SDK own fast HTTP retries, make outer retries slow and checkpointed, and cap background traffic so it can't starve live requests.
What does the Anthropic SDK retry by default?
The official Python SDK (anthropic) retries a failed request twice by default, so one client.messages.create() can become 3 HTTP requests. It retries connection errors, 408 Request Timeout, 409 Conflict, 429 Rate Limit and any 5xx, which includes the 529 overloaded error. It uses exponential backoff with jitter and respects the retry-after header when the server sends one.
This is a good default. It's also invisible. Nothing in your code says "retry," so you forget it exists, and then you write your own retry on top.
You can change it per client or per call:
import anthropic
client = anthropic.Anthropic(max_retries=0) # whole client
client.with_options(max_retries=5).messages.create(...) # one call
Where did the 36 come from?
The system is Preterview, an interview practice tool I built and run (full disclosure: I built it). A user does a realistic voice interview, and when the session ends, a background job sends the transcript to Claude to score it and write a report. You can see it at preterview.com/en. The live interview turns and the background report jobs shared one API key, which matters later.
Here is roughly what the report path looked like:
import anthropic
from tenacity import retry, stop_after_attempt, wait_fixed
from rq import Retry
client = anthropic.Anthropic() # max_retries=2, I never typed it
@retry(stop=stop_after_attempt(4), wait=wait_fixed(2))
def call_claude(**kw):
return client.messages.create(**kw)
def build_report(session_id):
rubric = call_claude(...) # step 1: score against rubric
evidence = call_claude(...) # step 2: pull quotes per criterion
report = call_claude(...) # step 3: write the report
save(session_id, report)
queue.enqueue(build_report, session_id, retry=Retry(max=2))
Three layers, three people's worth of good intentions:
| Layer | Attempts | Who added it |
|---|---|---|
| Anthropic SDK | 3 | The SDK, silently |
| tenacity | 4 | Me, "just in case" |
| RQ job retry | 3 | Me, months earlier |
| Product | 36 | Nobody |
And there was a fourth multiplier hiding in build_report. An RQ retry restarts the whole function. If step 3 failed, the retry re-ran steps 1 and 2, which had already succeeded. My worst job that evening made 42 requests and still failed.
Why did a retry storm break live interviews?
Background retries piled up during the outage and all fired at once when the API recovered, which pushed my input tokens per minute past my rate limit. The live interview turns, which used the same key, started getting 429s.
The timeline, from my logs:
-
Minute 0 to 22: intermittent 529s. Eight report workers spent most of the window sleeping in backoff, holding their slots.
wait_fixed(2)in tenacity has no jitter, so workers that failed together retried together. - Minute 22: overload clears. Every worker wakes up, plus the backlog of queued reports, plus RQ retries re-running steps that had already succeeded. Report prompts carry a full interview transcript, so each request is heavy on input tokens.
- Minutes 22 to 28: input-token rate limit exceeded. Live turns got 429s on 64 of 352 turns (18%). The voice layer fell back to a "give me a second" filler and retried, which users notice.
The numbers for the report worker in that window:
- 140 report jobs, which should have been about 420 requests (3 calls each)
-
1,302 calls to
call_claude()according to my app logs - 3,904 actual HTTP requests according to the httpx hook
- Slowest report delivered 19 minutes after the session ended (normal is about 40 seconds)
Nothing here was a bug in the SDK. The SDK did exactly what it documents. I just never did the multiplication.
How do you count real HTTP requests from the Anthropic SDK?
Pass your own httpx client with an event hook. Your application logs count function calls; this counts what actually goes over the wire, including the SDK's internal retries.
import httpx, anthropic
def on_request(request):
metrics.incr("anthropic.http_requests", tags={"path": request.url.path})
client = anthropic.Anthropic(
http_client=anthropic.DefaultHttpxClient(
event_hooks={"request": [on_request]}
)
)
Put the ratio of HTTP requests to app-level calls on a dashboard. Mine normally sits around 1.01. During the incident it hit 3.0, which is the SDK's ceiling and a clear sign every call was exhausting its retries.
How should you layer retries with the Anthropic SDK max_retries?
Decide which layer owns which timescale, then write the product of all layers in a comment next to the code. Here is what I changed.
1. The SDK owns fast HTTP retries. Nobody else does. I deleted tenacity around API calls entirely. The SDK already has jittered exponential backoff and reads retry-after, which my wait_fixed(2) ignored.
2. Separate clients per traffic class.
# Live voice turns: latency beats persistence. One retry, then filler.
live = anthropic.Anthropic(max_retries=1, timeout=20.0)
# Background reports: patient, but bounded.
batch = anthropic.Anthropic(max_retries=3)
3. Outer retries are slow and checkpointed. The job retry now waits minutes, not seconds, and each step's result is stored so a retry resumes instead of restarting.
def step(session_id, name, **kw):
if (hit := store.get(session_id, name)):
return hit
if breaker.is_open():
raise CoolingDown()
with input_token_bucket.reserve(estimate_tokens(kw)):
out = batch.messages.create(**kw)
store.put(session_id, name, out)
return out
# Retry product: 4 (SDK) x 3 (RQ) = 12 attempts max,
# and the RQ attempts are 3 and 10 minutes apart.
queue.enqueue(build_report, session_id,
retry=Retry(max=2, interval=[180, 600]))
4. A circuit breaker on 529s. If more than 30% of the last 20 background calls got 529, workers stop pulling report jobs for 60 seconds. Reports wait in the queue instead of hammering a struggling API.
5. Background traffic gets a token budget. A simple token bucket caps the report worker at half of my input-tokens-per-minute limit. The recovery burst can still happen, but it can't eat the half that live turns need.
Did the fix work?
A week later there was another overload stretch, about 15 minutes. The report worker sent 1.3Γ its normal request volume instead of 9.3Γ, the HTTP-to-call ratio peaked at 2.4, and live turns saw zero 429s. The slowest report arrived 11 minutes late, which is worse than 40 seconds but much better than a failed report with a 42-request bill attached.
The honest caveat: I have two incidents, not a controlled experiment. The second outage was shorter and may have been milder. What I can say for certain is the arithmetic: the old setup could reach 36 attempts per failing call by design, and the new one is capped at 12, spread over 13 minutes.
Checklist before your next outage
- Grep for
@retry,Retry(,backoff, andmax_retriesin any code path that touches an LLM client. Multiply what you find. - If you wrap the SDK in your own retry, set
max_retries=0on that client. Pick one. - Never use fixed waits without jitter for shared upstreams.
- Make multi-step jobs resumable before you make them retryable.
- Keep live and background traffic from competing for the same rate-limit headroom.
So, what does Anthropic SDK max_retries actually cost you?
By itself, very little: the Anthropic SDK retries each request 2 times by default with jittered backoff, which is a sensible setting. The cost appears when you stack it under your own retry decorator and a job-queue retry, because retry layers multiply. My 3 Γ 4 Γ 3 setup allowed 36 attempts per failing call, produced 3,904 requests where 420 were needed, and the recovery burst caused 429s on 18% of live turns. Let the SDK own fast HTTP retries, make outer retries slow and checkpointed, write the retry product in a comment, and budget background traffic so it can't starve the requests your users are waiting on.
Written by the developer behind Preterview, an interview prep platform.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes β full credit and traffic to the original publisher.