AML.T0054
MITRE ATLAS / LLM Jailbreak
Adversaries induce a language model to ignore, circumvent, or override its safety and alignment behavior in order to elicit output the model was intended to withhold. Techniques include role play, encoding, and multi-turn pressure.
In plain English
Guardrails are behavior, not walls. Enough framing, indirection, and persistence will usually convince a model that this particular request is the exception.