System Prompt Leak
Also known as:Prompt Extraction
System Prompt Leak: Disclosure of a language model’s hidden system instructions to the user. The leak is usually the first stage of a prompt injectionPrompt InjectionInjecting instructions into a language model’s input to override its intended directives. chain, because it exposes the very rules that need to be circumvented.
How it works and where it fits
The system prompt describes role, tone, permitted topics, and connected tools — and in poorly built applications also credentials, internal identifiers, or business logic. It is extracted through direct requests, through asking for a summary or translation of the preceding instructions, or by asking for it in encoded form. Partial answers are revealing too: merely describing a condition often shows whether a check is implemented as a hard comparison or as soft wording.
Practical security relevance
A system prompt cannot reliably be kept secret — it sits in the same context as the user input. From this follows the only dependable rule: no secrets in the prompt. Credentials, keys, and internal endpoints belong in a layer the model cannot read. Operationally the leak still constitutes a reportable finding, because it substantially enlarges the attack surface for subsequent bypasses.
Related concepts
- Prompt InjectionPrompt InjectionInjecting instructions into a language model’s input to override its intended directives.: Injecting instructions into a language model’s input to override its intended directives.
- LLM Penetration TestingLLM Penetration TestingAuthorized security testing of Large Language Model deployments for vulnerabilities such as prompt injection, system prompt extraction, RAG poisoning, and tool abuse.: Authorized security testing of Large Language Model deployments for vulnerabilities such as prompt injection, system prompt extraction, RAG poisoning, and tool abuse.
- Data ExfiltrationData ExfiltrationUnauthorized transfer or theft of data from an organization.: Unauthorized transfer or theft of data from an organization.
- AI Penetration TestingAI Penetration TestingAuthorized security testing of AI and machine-learning systems for vulnerabilities such as prompt injection, model extraction, and training data leakage.: Authorized security testing of AI and machine-learning systems for vulnerabilities such as prompt injection, model extraction, and training data leakage.