Guardrails

Guardrails are checks placed around a model to catch or block a problem before it reaches a user or causes harm — filtering unsafe output, blocking a disallowed action, validating a response before it's used. They sit outside the model itself, in application code the model doesn't control, which is exactly why they can enforce something a prompt instruction alone can't.

Two places they apply

Input guardrails check what's about to be sent to the model — blocking an obviously malicious request, stripping or flagging sensitive personal data before it leaves the application, refusing to process a category of request the system was never meant to handle.

Output guardrails check what the model just produced, before it reaches a user or triggers an action — content moderation, staying within an allowed topic, catching a response that doesn't match the required structured output schema, or flagging a claim that isn't actually backed by the source material it was supposed to answer from.

How they're actually built

A guardrail can be as simple as a rule or a pattern match — reject anything containing a specific term, require a field to be present. It can be a small classifier trained for one specific check, like detecting personal data. Or it can be another model call, scoring the output against a policy the same way LLM-as-judge scores it against a quality rubric — the mechanism is similar, but the judgment is "does this violate a policy," not "is this a good answer." Real systems usually stack more than one kind, since a cheap rule check catches the obvious cases before a more expensive model-based check has to run at all.

A layer, not a complete fix

Guardrails reduce risk; they don't eliminate it. A rule-based filter can be worded around, and a model-based check can be wrong in the same ways any model call can be wrong. The prompt injection page makes this same point about defenses generally: no single layer is a complete fix, which is why real systems combine several — guardrails alongside narrow tool permissions, human approval on sensitive actions, and treating untrusted content as untrusted in the first place — rather than treating any one of them as sufficient on its own.

When to use them

Any system that takes input from outside its own trusted boundary, or produces output that reaches a real user or triggers a real action, benefits from at least basic guardrails. How many, and how strict, should scale with the stakes: a low-risk internal tool needs less than a customer-facing system that can take real actions on someone's account.

In this guide
  1. Two places they apply
  2. How they're actually built
  3. A layer, not a complete fix
  4. When to use them
  5. FAQ

FAQ

Is schema validation on a structured output the same thing as a guardrail?

It's one kind of output guardrail — a narrow one that checks shape rather than policy or safety. A full guardrail setup usually includes shape validation alongside separate checks for content that's unsafe, off-topic, or unsupported by the source material, not shape validation alone.

If a system has guardrails in place, is prompt injection still a risk?

Yes. Guardrails can catch some injected instructions and some unsafe outputs, but they're not an enforced boundary the way a permission check is — a well-crafted injection can still get through a filter that wasn't built to catch that specific wording. Guardrails reduce the risk; they don't remove the need for the other defenses on the prompt injection page.

Practice interview questions on AI safety →