๐ŸงฐDaily Toolbox
โ† All guides
API

API Rate Limiting: Why Your 429 Error Isn't What You Think

2026-09-04 ยท 7 min read

Your payment processor is down. Logs show HTTP 429 everywhere. "We're sending 100 req/s, their limit is 1000/min. The math works."

The math is wrong.

Real APIs enforce three separate limits simultaneously. Violate any one โ†’ instant 429.

The 3 Rate Limits Nobody Tells You About

### 1. Per-Second Burst Limit (the silent killer)

Stripe's docs say "100 requests per second." You send 95 req/s. You get 429s.

Why? Stripe's actual limit is 100 req/s average over 10 seconds + 25 req/s maximum in any single second.

Send 95 requests in the first second? The per-second burst cap rejects you before the average limit even notices.

Real example from 2023: A crypto exchange sent payment webhooks in batches of 40 requests. Stripe's 25/s instantaneous cap rejected 15 per batch โ€” $180K/day in failed payments before they caught it.

### 2. Concurrent Connection Limit (the invisible wall)

OpenAI's API limits:
- 3,500 requests per minute (documented)
- 500 concurrent connections (NOT documented)

Your load test sends 1,000 requests with Promise.all(). Each takes 2 seconds. You're under 3,500/min.

Result: 500 succeed. 500 get 429.

Why? You had 1,000 in-flight requests at once. Once OpenAI's 500 connection slots fill, new requests are rejected even if you haven't hit the rate limit.

The fix: Connection pooling libraries like p-limit or Bottleneck cap concurrent requests.

### 3. Token Bucket vs. Fixed Window (the timing trap)

Fixed Window (what you think):
- 0:00-0:59 โ†’ 1000 requests allowed
- 1:00-1:59 โ†’ counter resets

Token Bucket (what actually happens at AWS, Cloudflare, Stripe):
- You have a bucket with 1000 tokens
- Tokens refill at 16.67/second (1000/min รท 60s)
- Each request consumes 1 token
- Empty bucket โ†’ 429, even if you "waited a minute"

Real incident (2025): A startup sent 5,000 Twilio SMS at 2 AM daily. Limit: 1,000/min. Split into 5 batches, 1 minute apart.

Batch 1 (2:00 AM): 1,000 sent โœ…
Batch 2 (2:01 AM): 1,000 sent โœ…
Batch 3 (2:02 AM): 400 sent, 600 rejected โŒ

Why? Between 2:00 and 2:01, the bucket only refilled 1,000 tokens. Batch 1 drained it. Batch 2 drained the refill. By 2:02, only 400 tokens had refilled.

How to Debug Your 429s

### Check the Response Headers

Every major API returns these (RFC 6585):

- X-RateLimit-Limit โ†’ Your quota
- X-RateLimit-Remaining โ†’ Tokens left right now
- X-RateLimit-Reset โ†’ Unix timestamp when bucket refills
- Retry-After โ†’ SECONDS until you can retry (not a timestamp!)

Common mistake: Treating Retry-After as a Unix timestamp instead of a duration in seconds.

### Log Request Timestamps

If you're getting 429s but your logs show only 200 requests in the last minute (limit is 1000), you're hitting a different limit โ€” burst or concurrency.

### Measure Time-to-First-Byte (TTFB)

If requests take 5 seconds each and you send 300/min, you have 1,500 concurrent connections at peak (5s ร— 300/min รท 60s).

High TTFB + 429s = you're hitting the concurrent connection limit, not the rate limit.

The 4 Strategies That Actually Work

### 1. Exponential Backoff with Jitter

Don't just retry after Retry-After seconds. Add random jitter (0-1000ms) to prevent thundering herd when many clients retry at the same time.

### 2. Request Queuing

Libraries like Bottleneck handle:
- Token bucket refills
- Concurrent connection limits
- Per-second burst caps

Configure: initial tokens, refill rate, max concurrent, min time between requests.

### 3. Priority Queues

Not all requests are equal. Payments > analytics.

Use separate queues with different concurrency limits. When you hit 429, low-priority requests wait longer. Critical flows keep running.

### 4. Circuit Breaker

If an API returns 429s for 30 seconds straight, stop sending requests. You're burning tokens and making it worse.

Circuit breakers open after a threshold (e.g., 50% failure rate), instantly return fallback responses for 30s, giving the API time to recover.

The Hidden Cost of 429s

Each rejected request costs:

1. Wasted API quota (some APIs count 429s against your limit)
2. Latency (retry delays add up)
3. Money (Stripe charges $0.05 per failed payment retry)
4. User trust ("Payment failed, try again")

Case study: A SaaS retried Stripe payments 5 times on 429. Each retry took 2 seconds โ†’ 10-second user-facing delay per failed payment.

After exponential backoff + circuit breaker: 0.8-second delay. Conversion rate up 3.2%.

Tools That Save You

1. [HTTP Header Analyzer](/tools/http-header-parser) โ€” Decode X-RateLimit headers
2. [JSON Formatter](/tools/json-formatter) โ€” Parse API error responses
3. Datadog API Monitoring โ€” Track 429 rates across endpoints

The One Rule to Remember

Rate limits are not a ceiling. They're a speed limit on a highway with three lanes, two toll booths, and a cop hiding behind a billboard.

You're managing:
- Burst caps (the cop)
- Concurrent slots (the toll booths)
- Token refill rates (the lanes merging)

Treat 429s as architectural feedback, not errors. Your system is saying: "Slow down, or redesign."

A 1000/min limit doesn't mean "send 1000 every minute." It means "stay under 1000 while the bucket refills, AND don't burst past 25/s, AND don't hold 500 connections open."

Most APIs won't tell you all three rules. Now you know where to look.

#API#Rate Limiting#Backend

Try the free tools mentioned above

Open dev tools โ†’