Handling OpenAI RPM, TPM, and TPD Limits at the Gateway Layer
Gateway-level rate limiting prevents cascading failures across teams sharing a single OpenAI quota.
Concentrate.ai
On this page
OpenAI enforces four separate rate limits on every API account: requests per minute, tokens per minute, requests per day, and tokens per day. Hitting the ceiling on any single one causes every request to get a 429, no matter how much headroom exists on the other three. That fact alone should change how a team architects its infrastructure, because handling this inside each individual service is a losing game. It belongs at the gateway, the shared layer every request already passes through.
LLM Gateway
Instant access to every model
Change models, providers, add fallbacks, and more by running all AI usage through one platform.
The windows aren't even fixed. RPM and TPM roll over a 60-second window, RPD and TPD over 24 hours. There's no midnight reset, no clean top-of-the-minute refresh. And OpenAI doesn't always enforce these limits smoothly either: an RPM of 600 might get quantized down to a hard cap of 10 requests at a time, so a short burst of traffic can trip a 429 even when the average across the full minute is comfortably under quota.
TPM adds another wrinkle worth knowing cold. It's counted against projected token usage, not just what the model actually returns, meaning configuration choices around token limits can cause quota to burn faster than real usage would suggest. Set max_tokens too high out of caution and quota burns even on requests that return a short response. That's quota lost to a configuration habit, not real usage.
How OpenAI's six-tier system determines your starting ceiling
Tiers move up automatically, based on how much an account has spent and how long it's been around. No support ticket required, at least not until an org outgrows Tier 5.
Tier 1 kicks in after a first payment of at least $5. GPT-4o at that tier runs 500 RPM and 30,000 TPM, which is the ceiling most teams are shipping their first production feature against. It sounds fine until real traffic causes the gaps to appear.
Tier 2 needs at least $50 in cumulative spend and 7 days since that first payment. GPT-4o jumps to 5,000 RPM and 450,000 TPM there, a substantial multiple over the previous TPM alone. And that 7-day clock cannot be rushed by throwing more money at OpenAI faster. Spend $50 in an hour and the tier still waits for the calendar to catch up.
Tier 3 asks for $100 cumulative spend plus the same 7-day minimum. RPM stays at 5,000, but TPM scales up substantially again. The pattern across tiers is consistent: token throughput scales much faster than request throughput, which matters a lot for teams sending large prompts or long context windows rather than firing off high volumes of small ones.
Individual services handling their own rate limit logic
The standard first move is exponential backoff with random jitter on a 429. Fine as a baseline. Not remotely sufficient once more than one service is calling OpenAI.
Each service backs off on its own, with zero visibility into what every other service is pulling from the same shared org-level quota pool at that exact moment. Service A slows down nicely. Service B, running on a different team's deploy schedule, has no idea A even exists. The org-level ceiling is one number, shared, regardless of internal team boundaries. It's one number, shared.
A detail that changes the math entirely: failed retries still count against rate limits. A retry storm, kicked off by three or four services simultaneously hitting backoff logic after a shared 429, doesn't relieve pressure on the quota. It adds to it. The fix people reach for actively makes the problem worse when it's deployed at the service level instead of centrally.
RPD and TPD compound the blindness. A service can space its requests out perfectly within every 60-second window, pass every per-minute check with room to spare, and still blow through the daily token limit with no warning before the 429 occurs. Per-minute backoff logic literally cannot see a daily ceiling coming.
The gateway as the architecturally correct layer for rate limit absorption
A gateway sits between the applications and the AI provider. It's already the natural chokepoint for security, governance, and observability, so absorbing rate limits there is just what a central request layer is for. It's just what a central request layer is for.
Because every request from every service passes through one point, the gateway can track rolling consumption across all four dimensions, RPM, TPM, RPD, TPD, in real time. It can make a routing call before a 429 ever gets generated, not after.
That capability breaks down into four concrete mechanisms:
Queuing and request shaping hold and space out requests to stay inside the rolling window, instead of firing everything at once and retrying on failure. Automatic failover kicks in when a provider returns a retryable error, 429, 500, 502, 503, or 504, and routes the request to a backup model or provider transparently. The calling service just sees a normal response. Weighted load distribution across multiple API keys or providers stops any single key from eating more than its share of the shared pool. And per-consumer quota enforcement lets the gateway apply its own inbound limits on internal teams, independent of whatever OpenAI enforces outbound, so one team's traffic spike doesn't starve everyone else.
Gateway-level rate limiting has to be token-aware, not just request-aware. Counting requests alone tells a team nothing about TPM exposure. Accounting for the actual token weight of each call is the baseline requirement before any of this works.
Queuing, fallback routing, and cross-provider load balancing in practice
Picture a request that would push TPM past the rolling ceiling. Instead of firing it off and eating the 429, the gateway holds it in a managed queue and releases it at a pace the provider will actually accept. No wasted retry, no compounding pressure on the shared pool.
Priority queuing sharpens that further. A user waiting on a chat response needs to be handled now, ahead of a background batch job summarizing yesterday's logs. A gateway that understands priority can let the interactive path stay fast even while quota is tight, and push the batch job back a few seconds.
Fallback routing is where this gets genuinely useful. When the primary model throws a 429 or a 5xx, the gateway can reroute to a secondary model, a lighter version, a different provider, without the calling service ever knowing a failure happened. That decision lives in policy configuration, not scattered across if-statements in a dozen codebases.
Cross-provider fallback takes it one step further: routing to Anthropic or Google when OpenAI capacity runs dry. A unified API abstraction underneath is what makes that work cleanly. Otherwise every fallback target needs its own bespoke integration code, and the whole point of centralizing this logic falls apart.
Inbound per-team quota enforcement and multi-tenant isolation
Because OpenAI enforces limits at the org level, one team's misbehaving pipeline, a loop that retries too aggressively, a job that fires far more requests than intended, can burn through quota that every other team also depends on. It's a governance gap that has to be closed structurally.
The fix is per-consumer limits enforced at the gateway itself: RPM and TPM ceilings applied to each internal team or tenant, throttling requests inbound before they ever touch the shared provider-side pool.
Two flavors of that policy matter, and they're not the same thing. Hard ceilings block a consumer once they hit their allocation, full stop. Soft budgets alert or reroute as a team approaches its limit, without cutting them off outright. Both are policy decisions, and both belong in the gateway's configuration, not buried in application code where nobody remembers they exist six months later.
Per-consumer tracking has a second payoff that's easy to underrate: it produces the exact data needed for chargeback, allocating token spend back to whichever team actually generated it. That's the same discipline that made cloud cost management work once organizations got serious about it, applied directly to AI spend.
What the gateway must expose to engineering, finance, and security teams
Real-time visibility into rolling RPM, TPM, RPD, and TPD consumption, broken down by team, key, model, and provider, has to be available as it happens. Not reconstructed weeks later from an invoice after the damage is already done.
Retry and fallback telemetry answers a different set of questions: which requests triggered a retry, which of the four dimensions caused it, which fallback provider picked up the slack, and what that did to latency. That's the exact signal an engineering team needs to actually tune routing policy instead of guessing.
Finance needs its own view: token spend broken down by team, application, environment, and model. The same granularity that made cloud compute budgeting workable applies here, and without it, AI spend becomes a mystery line item that appears on the bill at the end of the month.
One more piece gets overlooked constantly: prompt caching. Many teams haven't turned it on, and they're paying full price on input tokens for content that repeats across requests. A gateway that's already tracking token flow is positioned to surface exactly where caching would cut spend, rather than leaving every service team to notice that on their own, or miss it.
Evaluating gateway options: managed service versus self-hosted
Self-hosted gateways put the infrastructure in the engineering team's hands directly. LiteLLM is the most widely evaluated option in this category, built around a Python-centric interface, and it handles protocol translation along with cost-based routing. Its cost-based routing picks the cheapest deployment by price without quality optimization, though its Auto Router v2 goes further, optimizing for cost against quality using complexity-based, semantic, and adaptive routing. Even so, that's infrastructure-level routing logic, not a system reasoning about output quality in any deep sense.
The tradeoff that gets glossed over: self-hosting means the gateway itself becomes one more piece of infrastructure that has to be deployed, patched, scaled, and secured, by the same team that's supposed to be shipping product. That's a real, ongoing cost, and it belongs in any honest build-versus-buy comparison, not as a footnote.
The broader landscape includes other approaches, such as fully managed hosted platforms that abstract away provider access entirely and learned routing systems built specifically to optimize the cost-versus-quality tradeoff using trained models rather than static rules. Each fits a different operational profile, and picking one is less about which is "best" and more about which matches the team's appetite for running infrastructure.
On the managed side, Concentrate offers a unified API connecting to a wide range of LLM providers, OpenAI, Anthropic, and Google included, without requiring separate provider keys, custom integration code per provider, or self-hosted infrastructure to babysit. The operational weight of multi-provider rate limit handling, fallback routing, and quota management sits inside the platform itself rather than landing on an engineering team's sprint backlog. For teams weighing the true cost of self-hosting against the value of getting this handled elsewhere, that tradeoff deserves the same scrutiny as any other infrastructure decision.