Alok Upadhyay | June 2026 · Explainer

TL;DR: If your safety behaviour lives entirely in the weights, then every policy change is a training run, and the model’s refusal boundary is a snapshot of what someone believed months ago. Treating policy as retrievable data (fetch the clauses that apply, inline them as explicit constraints, check the output against the same clauses) decouples the two. It also creates a new attack surface, and pretending otherwise is how these systems fail.


The mismatch

Safety alignment bakes a policy into weights. That works, and for broad norms it is the right place for them: a model should not need a retrieved document to decline to help synthesise a nerve agent.

The trouble is the long tail. Real deployments accumulate policy that is specific, conditional and volatile:

  • what counts as financial advice in one jurisdiction but not another
  • which medical claims need a referral disclaimer
  • what an enterprise customer contractually forbids the assistant to discuss
  • a clause added last Tuesday because of an incident last Monday

None of that is stable enough to train on. By the time a fine-tune lands, some of it is wrong. And weights are opaque in the way that matters here: when the model refuses, you cannot point at the rule it applied, because there isn’t one, only a gradient-shaped disposition. For anything subject to audit, “the model declined, we think because of how it was trained” is not an answer.

Policy as data

The alternative is to stop treating policy as something the model is and start treating it as something the model is told, per request.

Inference-time architecture: risk router, policy retrieval, inlined constraints, and an output check reading the same clauses The weights never change. What changes is which policy text is in the context, and that is a data edit under review, not a training run.

Four moving parts:

1. Risk router. Cheap classification of the request into policy domains: medical, financial, legal, self-harm, tenant-specific. This is a routing decision, not a safety decision, and it should be biased toward over-triggering. Retrieving a clause that turns out to be irrelevant costs tokens. Missing one costs a violation.

2. Retrieval. Pull the clauses for those domains from a versioned store. Not a similarity search over a wiki; clauses are short, identified, and fetched by domain and jurisdiction. Embedding similarity is a reasonable recall mechanism but a poor authorization mechanism, and the distinction matters when a near-miss silently returns nothing.

3. Inlining. Place the retrieved clauses in the prompt as explicit constraints, clearly delimited, with their ids. The model now has the actual rule in front of it rather than a remembered impression of one.

4. Output check. Verify the response against the same clause set before it ships. Same source, second look: the generation step is persuadable in ways a targeted check is not.

What this buys is mostly operational, and the operational part is the point:

  • Auditability. The log says policy.fin.003 fired. You can show the clause text, its version, and who approved it.
  • Iteration speed. Editing a clause is a reviewed document change, live in minutes, revertible.
  • Variation. Per-jurisdiction and per-tenant policy without a model per tenant.
  • Testability. Clauses are discrete, so you can write a test per clause instead of probing a disposition.

Where it leaks

This is the part that gets skipped in architecture diagrams, so let me be blunt about it. Inlining policy means policy now lives in the context window, and the context window is shared with text an attacker controls.

Trust boundary: curated policy clauses in a trusted channel versus user text, documents and tool output in an untrusted channel Both channels occupy one context window. Keeping them distinct is a prompt-construction discipline; the model does not infer the boundary on its own.

Prompt injection against the policy block. The oldest problem, with a new target. Ignore the policy above is the naive version; the effective versions arrive inside a retrieved document or a tool result the user never typed. If your pipeline retrieves a web page and pastes it in next to your policy clauses, that page is now arguing with your policy from inside the same context. Instruction-tuned models are built to follow instructions, and they are not reliable at deciding which region of the prompt is allowed to issue them.

Retrieval miss as silent allow. If the router mislabels a request and nothing relevant is retrieved, the model answers with no constraint present, and the log shows a clean request with no clause fired. This is the failure I would worry about most, because it is indistinguishable from correct operation unless you design for it. A retrieval miss in a search product is a worse answer; a retrieval miss in a policy layer is an unguarded one. Empty retrieval in a risk-flagged domain should deny, not proceed.

Conflicting clauses. Two clauses, both retrieved, pointing opposite ways: a tenant permitting something a jurisdiction forbids. Absent an explicit precedence rule, the model picks, which means the resolution is whichever clause happened to sit lower in the prompt. Precedence belongs in the assembly step, decided by code.

Context budget. Policy competes with the user’s actual content. Retrieve too eagerly and you spend the window on clauses that do not apply, degrading the answer and tempting someone to trim the safety block first.

Over-refusal drift. Broadly worded clauses are cheap to write and expensive to live with. A clause that reads any discussion of medication will refuse a question about ibuprofen dosage on a label. Because clauses are easy to add and nobody is scored on removing them, the failure accretes.

Defences that actually help

separate the channels    policy in a delimited system region; user
                         content and tool output never in that region

deny on empty            no clause retrieved in a flagged domain -> refuse
                         and escalate; never silently proceed

precedence in code       resolve clause conflicts before the model call,
                         not by prompt ordering

check independently      validate output against clause ids with a
                         separate call; do not ask the generating model
                         to grade itself in the same turn

log the clause ids       an answer with no clause id in a flagged domain
                         is an incident, not a success

test per clause          every clause ships with a should-refuse and a
                         should-allow example; the second prevents drift

The should-allow case deserves emphasis. Safety suites are usually all attacks, which makes over-refusal invisible: a model that refuses everything scores perfectly. Pairing each clause with a benign example that must still be answered is what keeps the boundary honest.

What this is not

It is not a replacement for alignment training. Retrieved policy is a specificity layer on top of weights that already refuse the obvious. If the base model will help with the clearly catastrophic request whenever no clause is retrieved, the architecture is upside down: retrieval becomes the only thing standing between a routing bug and a serious failure.

Nor is it a guarantee. Every clause in the context is a clause an attacker can attempt to argue with, and every retrieval step is one that can return nothing. The honest claim is narrower and still worth a lot: it makes policy legible, versioned and fast to change, and it converts a class of silent failures into logged ones.

That trade, accepting a known attack surface in exchange for auditability and a deployment cycle measured in minutes, is usually the right one. It is only the right one if you instrument the failure modes rather than assuming the diagram holds.


The one-sentence version

Put the policy in the context, not only in the weights. Then treat that context as contested territory, because it is.