Understanding which infrastructure layer solves your actual LLM cost and control problem.
Concentrate.ai
On this page
LLM proxy, router, and gateway are three layers of infrastructure that stack on top of each other. They're three layers of infrastructure that stack on top of each other, and most vendors have muddied the terms so badly that picking the wrong one costs teams either months of wasted engineering effort or a compliance gap they don't find out about until an auditor does. LiteLLM calls itself both a proxy and a gateway. Helicone's own docs use the two words interchangeably. That's not a branding quirk, it's a real signal that the market hasn't settled on what these words mean, which leaves engineers guessing at what a tool actually does from its name alone.
LLM Gateway
Instant access to every model
Change models, providers, add fallbacks, and more by running all AI usage through one platform.
This matters because the layers solve different problems, and skipping one or piling on one you don't need yet both carry a cost. What follows breaks the stack apart, layer by layer, then walks through how to tell which one an actual team, at an actual stage of growth, actually needs.
The transport layer: what a proxy ignores
A proxy sits between an application and a model provider, catches the outgoing HTTP request, and forwards it along. From the application's side, almost nothing changes. Swap one base URL, and the rest of the code stays exactly as it was.
That simplicity is the whole point. A proxy has no idea who's calling it. It doesn't know if the request came from the marketing team or the fraud detection pipeline, whether that team has much of its monthly budget left, or whether this is request number ten or number ten thousand. It just forwards traffic.
A proxy built well for LLM traffic handles a specific, bounded set of jobs:
Forwards requests and pools connections so every call isn't opening a fresh socket
Terminates TLS
Swaps credentials, so the client sends its own key in and the proxy sends the provider's key out, meaning application code never touches the real upstream secret
Counts tokens and logs cost per request, at a basic level
Caches exact-match responses so the same query twice doesn't get billed twice
Passes streaming responses through without breaking them
What it doesn't do matters just as much. A proxy doesn't choose between models. It doesn't enforce a rate limit per team, track a budget, or apply any compliance rule. Those jobs belong to layers built above it.
A plain HTTP proxy has no idea what a 429 error from a provider actually means for an LLM request. It just sees a failed call. An LLM-aware proxy at least recognizes provider-specific error codes, even if it doesn't act on them, and that distinction separates a proxy from a router.
For early-stage projects, internal tools, and single-provider setups, a proxy is exactly the right amount of infrastructure. A proxy gets credential handling and basic caching out of the way without asking anyone to think about multi-tenant policy, which lets teams iterate faster at that stage than governance work would allow. Every layer above it depends on the proxy existing first. Its narrowness isn't a limitation to work around, it's the reason it's fast to build on.
The decision layer: how a router chooses which model handles each request
A router has one job: look at an incoming request and decide which model and which provider should handle it. The proxy then carries out that decision. Because a router is pure decision logic, with none of the transport plumbing a proxy handles, it can be tested on its own and swapped out without touching anything downstream of it.
Three routing strategies appear repeatedly in production systems:
Cost-based routing estimates how complex a request is and sends it to the cheapest model that can actually handle it. A simple classification task goes to a small, cheap model. Multi-step reasoning goes to a frontier model that costs more per token but doesn't fall apart on hard problems. Failover routing keeps a ranked list of providers and skips any one that a circuit breaker has flagged as down. This is a different job from cost optimization, even though the two share some of the same machinery, a candidate list and a policy. Treating them as the same thing is a common mistake, and it's the kind that causes production incidents: a router tuned only to save money won't necessarily fail over gracefully when a provider goes down. Metadata-based routing works off tags. The application sets a header, something like "support-bot" versus "code-review," and the router maps that tag straight to a model assignment. No complexity scoring involved.
The financial upside of doing this well is real and fast. Reporting from preto.ai puts typical cost reductions at 20 to 40 percent within the first week of turning on intelligent routing, which is a big enough number that most teams running multiple models should be asking why they haven't done it yet.
Evaluating routing strategies used to be a matter of vibes. RouterBench changed that: it offers more than 405,000 precomputed inference outputs across eleven LLMs and seven tasks, including MMLU, MT-Bench, MBPP, HellaSwag, WinoGrande, GSM8K, and ARC. That gives teams a real basis for comparing routing approaches instead of guessing.
Most routers optimize some aggregate utility signal and don't expose which individual factor drove a given decision. So a team can watch its bill go down and know routing is working, but diagnosing which specific signal started drifting in production is much harder, because the router isn't built to show its work at that level of detail.
LiteLLM's cost-based routing, for example, picks the cheapest available deployment. It's not optimizing for cost against quality, it's infrastructure-level routing, not a model that's learned what "good enough" looks like for a given task. That's a fine tool for the job it's built for. The difference matters for a workload where quality actually varies a lot between models.
The policy layer: what makes a gateway architecturally different from a proxy with routing bolted on
The thing that actually separates a gateway from a proxy with some routing logic tacked on is identity. A gateway knows which team, user, or application sent a given request, and it enforces rules based on that identity. A proxy has no concept of the caller at all. That's the real architectural line, not a longer feature list.
Once identity exists, a whole set of controls becomes possible that simply can't happen without it:
Multi-tenant auth, meaning virtual keys that map each caller to a specific tenant before the request ever reaches the model provider
A gateway functions as a control plane. Every request from every team, across every provider, passes through one enforcement point, which is what makes consistent governance possible without asking each engineering team to build its own version of budget tracking.
A gateway differs from an async observability tool like Langfuse, which sits outside the request path entirely and captures data after the fact. That approach adds zero latency to the actual request, but it can't enforce a budget in real time, can't cache a response, and can't route anything, because it's not in the path where those decisions happen. For a team that already has caching and routing handled elsewhere and just needs visibility, that's a deliberate and reasonable tradeoff, not a shortcoming.
How the three layers stack in practice
None of this is three separate products a team shops for independently. Proxy, router, and gateway are three layers of one stack, and most commercial products sold as "gateways" bundle all three, just at very different depths depending on the vendor.
LiteLLM's proxy layer is strong: over 100 supported providers behind one OpenAI-compatible API. Its router layer is present too, with multiple routing strategies including cost-based routing. The gateway layer is partial. Virtual keys and per-user budgets exist, but as noted above, the cost-based routing is infrastructure-level rather than something that's learned to weigh cost against quality.
Helicone's docs use "proxy" and "gateway" almost interchangeably, and the proxy layer is solid. The gateway-layer features work, but governance tools like RBAC, audit logs, and pre-inference budget enforcement are lighter than what enterprise-focused competitors offer.
Langfuse, again, is an async observer only. It never sits in the request path, so it adds no latency, but it also can't cache, route, or enforce a budget in real time, because those all require being in that path.
A gateway product deployed at the infrastructure level, running inside Kubernetes or a service mesh, gives strong control but comes with the operational overhead that kind of deployment always carries: someone has to run it, patch it, and scale it.
On the managed side, a service like Concentrate offers a unified API across multiple LLM providers, without requiring a separate key for each one, custom integration code per provider, or self-hosted infrastructure to keep alive. Proxy, router, and policy layers all ship bundled, and the model is built specifically against the idea of hidden per-token markups or a black-box router making decisions nobody can inspect, giving teams a visible, intentional say in where a given request actually goes.
The tradeoff between self-hosting and going managed is straightforward once it's spelled out. Self-hosted software like LiteLLM gives full control, but the engineering team now owns deployment, scaling, key rotation, and uptime. A managed gateway removes that operational load, but it means trusting a third party to run the control plane correctly.
The naming chaos has one practical fix: stop reading product names and start asking which layers a tool actually implements, and how deep each one goes. That question gets a real answer. The label on the pricing page doesn't.
Which layer a team actually needs depends on where they are operationally, not how ambitious they are
The right layer to adopt tracks operational reality, not ambition. A team with big plans but no traffic yet doesn't need a gateway. A team drowning in shared-key chaos across five departments needs one today, regardless of how small the company is.
A rough map by stage:
One team, one provider, low spend, no compliance obligations. Direct SDK calls are fine. Cost visibility starts to matter once real traffic reaches the system, which is when a proxy should be added. Multiple providers or models, cost becoming a real line item, still one team. Proxy plus router. The cost reductions from enabling model routing tend to pay for the added complexity almost immediately. Multiple teams sharing LLM access, compliance requirements starting to appear, spend becoming something several people are accountable for. The full gateway layer becomes necessary. Without identity-aware enforcement, per-team budgets and audit logs simply can't be built consistently, no matter how much custom middleware gets written to fake it. Enterprise, regulated industry, sensitive data flowing through prompts, finance or security asking for usage reports. Gateway-level PII redaction, RBAC, and structured audit logs stop being nice-to-haves and become baseline requirements.
The clearest sign a team has outgrown its proxy: engineers start hand-rolling middleware to enforce per-team limits, a security lead asks for an audit log the proxy has no way to produce, or a single shared provider key makes it impossible to say which project actually spent what.
Spending too little time on gateway infrastructure before there's meaningful traffic isn't the only risk; teams can just as easily over-invest in gateway infrastructure before there's any traffic to justify it. Spending days configuring gateway infrastructure before there's any meaningful traffic to govern is wasted effort, and a managed gateway sidesteps that entirely by making the policy layer available from day one, without a setup project attached to it.
Agentic workloads change this math faster than most teams expect. Analysis from Gartner and Atlan's LLM Cost Management for Enterprise 2026 found agentic workflows burn through 5 to 30 times more tokens per task than a standard chat interaction, and multi-agent setups can blow past initial cost projections by 3 to 10 times. A team building agent pipelines needs router and gateway infrastructure well before a team just shipping a chat feature does, because the cost and governance problems arrive faster.
The direction of travel across the industry backs this up. The gateway layer is moving from something only enterprises bother with to standard practice for anyone running more than one model in production.
The security and compliance obligations that only the gateway layer can fulfill
The exposure here is bigger than most teams assume. Cyberhaven's 2026 AI Adoption and Risk Report puts the share of enterprise AI interactions involving sensitive data at 39.7 percent, and a large chunk of that traffic runs through personal accounts that never touch corporate controls at all: 58.2 percent for Claude, 32.3 percent for ChatGPT.
Provider data retention terms compound the problem. OpenAI holds API data for 30 days for abuse monitoring. Anthropic cut its standard API log retention from 30 days down to 7 days in September 2025. Zero-data-retention is available, but only through a negotiated enterprise agreement, not on any pay-as-you-go plan. So the default, for most teams, is that sensitive data sits with a third party for at least a week, sometimes a month, without a contract saying otherwise.
OWASP's 2025 Top Ten moved Sensitive Information Disclosure up to position LLM02, a jump from position 6 back in 2023. The reasoning behind the move is straightforward: LLMs need more access to internal company data to be genuinely useful, and that access widens the surface for something sensitive to leak out.
Gateway-level PII redaction is the practical fix, and it works because it sits in the request path. It catches sensitive values before a request ever leaves the network, so no individual application has to build its own sanitization logic, and the whole redaction policy is reviewable from one place instead of scattered across a dozen codebases.
Masking is irreversible, which puts the resulting data outside GDPR's scope entirely, while tokenization is reversible through a vault, so the data is pseudonymized but still fully regulated. Which one a team should use depends entirely on whether the original value needs to come back later.
Regulatory timelines are already moving. The EU AI Act's prohibited-practices rules and AI-literacy obligations have applied since February 2025. General-purpose AI model obligations kicked in this past August. High-risk system obligations, covering risk management, data governance, record-keeping, transparency, and human oversight, land on December 2, 2027, after subsequent regulatory developments pushed the compliance date back. A widely referenced risk management framework for AI, under its Manage function, expects organizations to maintain audit records covering the key dimensions of each AI decision, all in service of incident response. HIPAA carries similar logging expectations for protected health information, with its own documentation retention requirements.
None of that is achievable at the proxy layer. A proxy has no identity model, keeps no per-team audit trail, and does nothing to redact data in-path. Meeting these obligations isn't a matter of configuring a proxy more aggressively, it requires the gateway layer, by definition, because the gateway is the only layer built to know who's asking and log what happened as a result.