Gateway-Layer PII Redaction Before Model Provider Submission
Redact sensitive data at the gateway, not downstream, to meet regulatory requirements.
Concentrate.ai
On this page
Sensitive data leaking into large language model requests is not a hypothetical risk to plan around someday. It is happening on a measurable scale right now, and the most defensible approach is stopping personal data before it ever leaves your own infrastructure. That means the gateway layer, the single choke point every outbound model call has to pass through, is where redaction belongs. Not in application code, not in a provider's post-hoc filter. At the door.
What the regulatory frameworks require, and where provider-side filters fail the test
Start with GDPR, because it's blunt about this. Sending personal data to an LLM provider counts as processing, full stop. Cumulative GDPR fines have exceeded €5.88 billion since 2018, and the European Data Protection Board's 2025 guidance addressed a common assumption teams were quietly counting on: pseudonymized data is still personal data if it can be re-linked to a person. Swapping a name for a token does not take the data outside the regulation's reach. You've just made the file slightly harder to read.
HIPAA works the same way, just with a sharper edge. Protected health information that shows up in a prompt sent to a provider raises disclosure concerns under HIPAA regardless of any filtering the provider applies downstream. It doesn't matter if the provider has some filter running downstream. The question regulators ask isn't whether a filter caught it eventually, it's whether the sensitive bytes crossed the wire. Once they did, the disclosure already happened.
LLM Gateway
Instant access to every model
Change models, providers, add fallbacks, and more by running all AI usage through one platform.
PCI DSS adds another wrinkle for anyone touching payment data: cardholder data reaching an external service pulls that service straight into your compliance scope. That's not a technicality, that's an audit finding waiting to happen.
Cisco's 2025 Data Privacy Benchmark Study found 64% of privacy and security professionals worry specifically about AI tools exposing company or employee data. That number isn't paranoia, it's people who read the regulations tracking the actual exposure.
Provider-side filters, the ones vendors advertise as "we don't train on your data" or "PII scrubbing included," run after the request has already crossed your perimeter. They can check a box. They cannot shrink your compliance boundary, because the sensitive data already left the building by the time any filter touches it. The EDPB's 2025 guidance makes the same point about pseudonymization: replacing a name with something like PERSON_001 reduces exposure, sure, but it doesn't move the workflow outside GDPR. Redaction isn't anonymization. A rare job title, an exact event date, a small-town location, that combination can still point straight at one person even with the name gone.
So the compliance issue collapses into a location issue. Redaction either happens inside your perimeter, before anything leaves, or it doesn't count for much.
The four places PII enters an LLM request, and why application-level scrubbing misses most of them
Engineering teams tend to picture PII exposure as a single moment, but a typed social security number is one door out of four, and it's the only one the calling code can actually see.
Direct user input is the obvious case. Someone pastes an email address, an account number, a phone number into the interface. Visible, inspectable, the easiest to catch, which is probably why it's the one most teams build a filter for and stop there.
Retrieved context is where things get harder. RAG pipelines pull documents into the prompt at query time, and those documents often contain PII nobody flagged when they were indexed. The knowledge base becomes the leak nobody thought to check, because the user never typed that data. It was already sitting in a document, waiting to get pulled into a prompt.
Conversation history compounds the problem over time. Multi-turn sessions accumulate sensitive values from earlier messages and replay them on every subsequent call, including turns the application has long since stopped actively inspecting. A name mentioned in turn two is still riding along in turn fourteen.
Logs and traces are the quiet one. Raw prompts written to observability backends can sit there indefinitely, well past the life of the actual request, even when the original model call was perfectly fine to send.
Getting all four right means every engineering team building on top of the model has to implement the same scrubbing logic, correctly, every time. One team skips the retrieval step, one team forgets logs, and now there's a leak somewhere in the stack that nobody notices until an audit or a breach forces the question.
Agentic systems make this worse, not better. A single user turn can fan out into a dozen model calls: tool selection, tool outputs fed back in as context, retrieval lookups, intermediate reasoning steps, a final synthesis call. Redacting just the first message a user sends covers almost none of that surface. Every outbound hop is its own exposure event.
The fix that actually works follows a sequence: minimize the data at the source by only pulling fields the task needs, detect in layers using both schema-aware rules and context-aware NER, then apply policy per entity type and purpose. Detection after the prompt has already been assembled is too late. By then the sensitive value is already baked into a payload heading out the door.
That's the argument for a single upstream enforcement point instead of trusting a dozen teams to get application-level discipline right every single time.
How gateway-layer redaction works as an architectural pattern
A gateway redaction layer is a proxy sitting between your applications and the model providers. Every outbound request passes through it, gets scanned, gets rewritten if it contains sensitive values, and only the sanitized version continues on to OpenAI, Anthropic, or wherever it's headed. The provider never sees the original.
The architectural case is simple: one enforcement point governs traffic to every provider your organization uses, whether that's OpenAI, Anthropic, AWS Bedrock, or Google Vertex AI. A single configuration covers sensitive data handling everywhere, and it doesn't depend on which application team wrote the code that generated the prompt.
The pipeline runs in three stages.
Detection identifies sensitive entities in the payload: direct identifiers like names, emails, phone numbers, government IDs; financial identifiers like card numbers and account references; network identifiers like IP addresses and device IDs; and credentials like API keys, tokens, and private keys.
Redaction action is a policy decision with three real options. Masking replaces a match with a typed token, something like [EMAIL_REDACTED], and works when the task can proceed fine without the original value. Blocking rejects the request outright, appropriate when substitution isn't safe, when policy forbids processing that data class at all, or when the scanner finds a credential sitting in the payload. Monitor mode just logs what would have matched without touching the traffic at all, useful for tuning detection rules but not, under any reasonable definition, a privacy control. A monitored data path is not a protected one.
Rehydration closes the loop for applications that need a coherent answer back. The provider only ever saw tokens, but the gateway can restore the original values before the response reaches the calling application, so the end user gets something usable.
Logs and trace exports need the same treatment as the live request, as a second, separate payload. Handling the outbound call but leaving raw prompts sitting in an observability backend just relocates the exposure instead of closing it.
LangSmith's LLM Gateway, in public beta as of late 2026, illustrates the rehydration mechanic clearly: when a redaction policy is active and trace capture is turned on, the trace reflects what the provider actually saw, meaning the redacted placeholders, not the original values. Rehydration happens before the response goes back to the caller, so traces recorded outside the gateway itself can still end up with restored values if that boundary isn't managed carefully.
Output redaction solves a different problem: it runs after the provider has already responded. It runs after the provider has already responded, and it can mask sensitive values the model generated or echoed back. But by the time output redaction runs, the provider has already processed the original request. It cannot protect the upstream boundary. It's a downstream cleanup step, not a substitute for catching things on the way out.
One operational detail matters more than it sounds like it should: what happens when the scanner itself breaks. If a PII or secrets detector is unreachable, slow, or throwing errors, the gateway should fail closed, blocking the request rather than letting an unscanned payload slip through. Fail-open defeats the entire point of the layer.
Detection tooling: available engines and where each fits
Two detection approaches exist, and they trade off differently.
Rule-based detection, regex and checksum validation, is fast and has a low false-positive rate for structured identifiers: credit card numbers, government-issued personal identification numbers, IBANs with a valid checksum. It's the right tool for the fast path, where latency matters and the pattern is well-defined.
Model-based detection, named entity recognition, is slower but necessary for anything that doesn't follow a fixed pattern: names, addresses, nationality or political affiliation buried in ordinary prose. You can't regex your way to catching "Maria Delgado lives near the old mill in Cortez" as a location and a name. That needs a model.
LangSmith's LLM Gateway, in public beta as of September 2026, pairs Presidio for model-based categories (names, locations, nationality, religious and political groups) with regex for rule-based categories (emails, US phone numbers, US SSNs). Its secrets detection covers a specific list: LangSmith keys, AWS, GitHub, GitLab, OpenAI, Anthropic, GCP, Azure, Slack, Stripe, PyPI, npm, and cryptographic private keys. The detection scope is intentionally narrow, anchored to recognizable token shapes so ordinary high-entropy text doesn't accidentally trip a redaction rule.
Presidio itself, originally built at Microsoft and moving toward the independent Data Privacy Stack organization in 2026, handles both named entities and pattern-based structured identifiers, and it's the NER engine sitting inside LangSmith's gateway.
OpenAI's Privacy Filter, released April 2026 under Apache 2.0, takes a different approach: a local bidirectional token classifier with 1.5 billion total parameters and 50 million active parameters, running on a 128,000-token context window. It detects eight categories (private persons, addresses, emails, phones, URLs, dates, account numbers, secrets) and runs locally, so the prompt never gets shipped to a third party just to figure out whether it contains PII.
GLiNER2-PII, an open-source option, covers 42 entity types across seven categories including personal, contact, governmental, financial, digital identity, credentials, and dates. It's a useful reference point for what a mature production policy should probably aim to cover, even if a given deployment doesn't use every entity type.
AISIX's AI Gateway ships built-in detectors for emails, mainland China mobile numbers and resident ID numbers, bank cards, US SSNs, IP addresses, API keys, JWTs, and PEM private keys. Its structured number detectors cover a range of formats including bank cards, SSNs, and IP addresses.
None of these engines will catch an organization's own internal identifiers out of the box. A medical record number in MRN-format, an internal case ID, whatever homegrown scheme a company uses, that needs a custom rule, and every custom rule needs an owner and a real test suite: positive matches, near-misses, Unicode edge cases, and checks for how it interacts with other detectors running in the same pipeline.
Scan timeout is a policy lever worth taking seriously. LangSmith lets teams configure max processing time per policy, defaulting to 2 seconds with a range from 0.1 to 30, plus a timeout action set to either allow or block. Teams operating under strict compliance obligations should default that action to block. Letting a slow scan just wave a payload through defeats the purpose of having the scan.
What gateway-layer redaction does not cover
Redaction is not the same thing as anonymization, and conflating the two is where a lot of otherwise careful setups go wrong. Even a complete string removal doesn't guarantee anonymity. A rare job title, an exact event date, a small-town location, stack those three details together and you've re-identified someone even with every name stripped out. The right unit to assess isn't the single field, it's the full context: the prompt, the metadata, any attachments, and the model's output taken together.
Output redaction cannot undo an exposure that already happened upstream. If a model generates or repeats sensitive data in its response, output-side inspection can mask it before a human sees it, but the provider already processed the original request by that point. An output-only policy protects the end user from seeing the leak. It does nothing for the compliance boundary that already got crossed.
Streaming responses raise a separate cost. Inspecting a stream for sensitive content means buffering it, which pushes out time to first token and eats more memory. Teams need to weigh that latency hit against how strict the actual compliance requirement is before turning output redaction on for streaming endpoints.
Some content is simply outside the gateway's reach. Encrypted fields, images, audio, unsupported binary frames, none of that can be scanned by a text-based redaction engine. The gateway can only redact what it can parse. Opaque payloads need their own separate controls.
Monitor mode deserves repeating: it is not a privacy control. Matched content keeps flowing through the request path unaltered. It's a measurement tool for tuning detection thresholds before flipping enforcement on, nothing more. Calling a monitored path "protected" is a mistake that appears in an audit, not during normal operation.
Redaction also doesn't replace the rest of the security stack: access control, data classification, retrieval authorization, secure application design. It's one enforcement layer, and it works alongside those controls, not instead of them.
For agentic systems specifically, redaction needs to sit at every hop, not just the initial user message. It has to be a hard dependency in the request path for every outbound model call, not an optional piece of middleware that only fires on the first turn. Teams running agents should specifically audit whether their gateway enforces this for tool results and intermediate context, because that's usually where the coverage quietly stops.
False positives are a real cost too. Model-based detectors misclassify prose sometimes. Rule-based detectors over-match high-entropy strings that just happen to look like a credential. Before flipping any detector from monitor mode into enforcement, run it against real, representative traffic first and actually measure the false-positive rate. Guessing gets expensive fast, either in blocked legitimate requests or in redacted data nobody meant to touch.
Comparing implementations: LangSmith, Occludra, AISIX, Cloudflare AI Gateway, and Philter
LangSmith LLM Gateway, built by LangChain and in private beta as of September 2026, treats PII detection and secrets detection as independent toggles. PII detection requires entitlement enablement and is off by default for most organizations, but once entitled, it starts on by default for new policies. Secrets detection, by contrast, ships off by default regardless of entitlement status. Policy scope can be set at the organization, workspace, user, or API key level, and applies immediately once configured, no rollout delay. Redacted placeholders show up in gateway traces, reflecting what the provider saw, while the calling application gets rehydrated values back in its response. Max processing time is configurable between 0.1 and 30 seconds, with the timeout action defaulting to allow (configurable to block), and data retention plus data protection settings live together in a single data policy object. Organizations without the PII entitlement can request it through a web form for Enterprise customers or by reaching out to their account team, though availability during the beta remains limited.
Beyond LangSmith, the market includes gateway products from Occludra, AISIX, Cloudflare (AI Gateway), and Philter, each approaching the same core problem, intercepting and redacting before a provider ever sees the payload, with different tradeoffs in detector coverage, deployment model, and how deeply they integrate with logging and trace infrastructure. AISIX's built-in detector set, covering everything from mainland China identity documents to PEM private keys, along with structured number formats including bank cards, SSNs, and IP addresses, makes it a reasonable fit for organizations with international data footprints that need coverage beyond identifiers tied to a single country. Teams evaluating any of these should weigh the same questions raised throughout: where redaction happens relative to logging, whether timeout behavior fails open or closed, and whether the detector set matches the actual data types moving through the organization's own model traffic, rather than a generic list built for someone else's use case.
The sources checked for this guide are listed below.