AI Safety
The practical end of AI safety: how applications built on models actually get attacked, and how to limit the damage when something gets through.
- Guardrails — are checks placed around a model to catch or block bad input or output before it reaches a user, sitting outside the model itself.
- Human-in-the-Loop — means a person reviews or approves a specific action before it happens, rather than the system acting entirely on its own.
- Prompt Injection — is untrusted text getting read as instructions.