Prompt Injection

Prompt Injection: Injecting instructions into a language model’s input to override its intended directives. It is the central vulnerability class in LLM penetration testingLLM Penetration TestingAuthorized security testing of Large Language Model deployments for vulnerabilities such as prompt injection, system prompt extraction, RAG poisoning, and tool abuse. and heads the OWASP list for LLM applications.

How it works and where it fits

A language model draws no structural distinction between system instructions and user input — both are text in the same context window. Whoever controls the input can therefore place competing instructions. In direct prompt injection the user does so themselves. In the indirect variant the text comes from a source the model processes: a web page, a document, an email, a tool result. In agentic systems with tool access this turns into genuine attacker capability.

Practical security relevance

Output-side filters fall short because content can slip past them in transformed form — Base64, reversed, in another language. What works is architecture: secrets do not belong in the context of a component that processes attacker-controlled text; tool calls need their own authorization and confirmation for side effects; and content from external sources must consistently be treated as data rather than instructions.

  • LLM Penetration TestingLLM Penetration TestingAuthorized security testing of Large Language Model deployments for vulnerabilities such as prompt injection, system prompt extraction, RAG poisoning, and tool abuse.: Authorized security testing of Large Language Model deployments for vulnerabilities such as prompt injection, system prompt extraction, RAG poisoning, and tool abuse.
  • Injection AttackInjection AttackManipulates interpreters or applications via injected commands or data.: Manipulates interpreters or applications via injected commands or data.
  • Input ValidationInput ValidationVerification of input data regarding format, length, type, value range, and validity.: Verification of input data regarding format, length, type, value range, and validity.
  • Adversarial Machine LearningAdversarial Machine LearningDiscipline concerning the manipulation, deception, and securing of machine learning models.: Discipline concerning the manipulation, deception, and securing of machine learning models.