Your payment processor is down. Logs show HTTP 429 everywhere. "We're sending 100 req/s, their limit is 1000/min. The math works."
The math is wrong.
Real APIs enforce three separate limits simultaneously. Violate any one โ instant 429.
The 3 Rate Limits Nobody Tells You About
### 1. Per-Second Burst Limit (the silent killer)
Stripe's docs say "100 requests per second." You send 95 req/s. You get 429s.
Why? Stripe's actual limit is 100 req/s average over 10 seconds + 25 req/s maximum in any single second.
Send 95 requests in the first second? The per-second burst cap rejects you before the average limit even notices.
Real example from 2023: A crypto exchange sent payment webhooks in batches of 40 requests. Stripe's 25/s instantaneous cap rejected 15 per batch โ $180K/day in failed payments before they caught it.
### 2. Concurrent Connection Limit (the invisible wall)
OpenAI's API limits:
- 3,500 requests per minute (documented)
- 500 concurrent connections (NOT documented)
Your load test sends 1,000 requests with Promise.all(). Each takes 2 seconds. You're under 3,500/min.
Result: 500 succeed. 500 get 429.
Why? You had 1,000 in-flight requests at once. Once OpenAI's 500 connection slots fill, new requests are rejected even if you haven't hit the rate limit.
The fix: Connection pooling libraries like p-limit or Bottleneck cap concurrent requests.
### 3. Token Bucket vs. Fixed Window (the timing trap)
Fixed Window (what you think):
- 0:00-0:59 โ 1000 requests allowed
- 1:00-1:59 โ counter resets
Token Bucket (what actually happens at AWS, Cloudflare, Stripe):
- You have a bucket with 1000 tokens
- Tokens refill at 16.67/second (1000/min รท 60s)
- Each request consumes 1 token
- Empty bucket โ 429, even if you "waited a minute"
Real incident (2025): A startup sent 5,000 Twilio SMS at 2 AM daily. Limit: 1,000/min. Split into 5 batches, 1 minute apart.
Batch 1 (2:00 AM): 1,000 sent โ
Batch 2 (2:01 AM): 1,000 sent โ
Batch 3 (2:02 AM): 400 sent, 600 rejected โ
Why? Between 2:00 and 2:01, the bucket only refilled 1,000 tokens. Batch 1 drained it. Batch 2 drained the refill. By 2:02, only 400 tokens had refilled.
How to Debug Your 429s
### Check the Response Headers
Every major API returns these (RFC 6585):
- X-RateLimit-Limit โ Your quota
- X-RateLimit-Remaining โ Tokens left right now
- X-RateLimit-Reset โ Unix timestamp when bucket refills
- Retry-After โ SECONDS until you can retry (not a timestamp!)
Common mistake: Treating Retry-After as a Unix timestamp instead of a duration in seconds.
### Log Request Timestamps
If you're getting 429s but your logs show only 200 requests in the last minute (limit is 1000), you're hitting a different limit โ burst or concurrency.
### Measure Time-to-First-Byte (TTFB)
If requests take 5 seconds each and you send 300/min, you have 1,500 concurrent connections at peak (5s ร 300/min รท 60s).
High TTFB + 429s = you're hitting the concurrent connection limit, not the rate limit.
The 4 Strategies That Actually Work
### 1. Exponential Backoff with Jitter
Don't just retry after Retry-After seconds. Add random jitter (0-1000ms) to prevent thundering herd when many clients retry at the same time.
### 2. Request Queuing
Libraries like Bottleneck handle:
- Token bucket refills
- Concurrent connection limits
- Per-second burst caps
Configure: initial tokens, refill rate, max concurrent, min time between requests.
### 3. Priority Queues
Not all requests are equal. Payments > analytics.
Use separate queues with different concurrency limits. When you hit 429, low-priority requests wait longer. Critical flows keep running.
### 4. Circuit Breaker
If an API returns 429s for 30 seconds straight, stop sending requests. You're burning tokens and making it worse.
Circuit breakers open after a threshold (e.g., 50% failure rate), instantly return fallback responses for 30s, giving the API time to recover.
The Hidden Cost of 429s
Each rejected request costs:
1. Wasted API quota (some APIs count 429s against your limit)
2. Latency (retry delays add up)
3. Money (Stripe charges $0.05 per failed payment retry)
4. User trust ("Payment failed, try again")
Case study: A SaaS retried Stripe payments 5 times on 429. Each retry took 2 seconds โ 10-second user-facing delay per failed payment.
After exponential backoff + circuit breaker: 0.8-second delay. Conversion rate up 3.2%.
Tools That Save You
1. [HTTP Header Analyzer](/tools/http-header-parser) โ Decode X-RateLimit headers
2. [JSON Formatter](/tools/json-formatter) โ Parse API error responses
3. Datadog API Monitoring โ Track 429 rates across endpoints
The One Rule to Remember
Rate limits are not a ceiling. They're a speed limit on a highway with three lanes, two toll booths, and a cop hiding behind a billboard.
You're managing:
- Burst caps (the cop)
- Concurrent slots (the toll booths)
- Token refill rates (the lanes merging)
Treat 429s as architectural feedback, not errors. Your system is saying: "Slow down, or redesign."
A 1000/min limit doesn't mean "send 1000 every minute." It means "stay under 1000 while the bucket refills, AND don't burst past 25/s, AND don't hold 500 connections open."
Most APIs won't tell you all three rules. Now you know where to look.