LLM Jailbreak
Also known as:Jailbreak · Guardrail Bypass
LLM Jailbreak: Circumventing a language model’s safety policies through carefully crafted inputs. The term overlaps with prompt injectionPrompt InjectionInjecting instructions into a language model’s input to override its intended directives. but targets the model’s policies rather than the application’s instructions.
How it works and where it fits
Typical patterns are role-play in which the model is asked to portray a fictional character without limits, hypothetical framing, gradual escalation across several innocuous messages, and encodings or foreign languages that evade detection patterns. Automatically generated token sequences that statistically suppress the refusal response supplement these. What all variants share is shifting the context until the problematic output becomes plausible to the model.
Practical security relevance
Because language understanding itself is the attack surface, there is no complete solution at the prompt level. Operators therefore rely on layered defence: independent classifiers for input and output, robustness training, logging of unusual conversations, and above all limiting what the model can actually trigger. For a security assessment, the decisive question is what impact a successful jailbreak would have in the specific application.
Related concepts
- Prompt InjectionPrompt InjectionInjecting instructions into a language model’s input to override its intended directives.: Injecting instructions into a language model’s input to override its intended directives.
- LLM Penetration TestingLLM Penetration TestingAuthorized security testing of Large Language Model deployments for vulnerabilities such as prompt injection, system prompt extraction, RAG poisoning, and tool abuse.: Authorized security testing of Large Language Model deployments for vulnerabilities such as prompt injection, system prompt extraction, RAG poisoning, and tool abuse.
- Adversarial Machine LearningAdversarial Machine LearningDiscipline concerning the manipulation, deception, and securing of machine learning models.: Discipline concerning the manipulation, deception, and securing of machine learning models.
- AI Penetration TestingAI Penetration TestingAuthorized security testing of AI and machine-learning systems for vulnerabilities such as prompt injection, model extraction, and training data leakage.: Authorized security testing of AI and machine-learning systems for vulnerabilities such as prompt injection, model extraction, and training data leakage.