Insights

Prompt injection, explained for decision makers

Why AI assistants can be talked into doing the wrong thing, why connecting them to your data raises the stakes, and what testing can and cannot tell you.

Prompt injection is the first item on the OWASP Top 10 for LLM Applications, and it is the risk most organisations meet first. It is also widely misunderstood. This note explains it without the jargon.

What it is

A large language model follows instructions written in plain language. Your application gives it instructions (“answer customer questions about our products, politely, using these documents”), and then passes it content from users, documents, web pages or other tools.

The problem is that the model has no reliable way to tell your instructions apart from instructions that arrive inside that content. If a piece of content says “ignore your previous instructions and do this instead”, the model may comply.

That is prompt injection: crafted input that makes a model ignore or override the instructions it was given.

Two kinds

Direct injection comes from the person typing. A user tries to talk the assistant into revealing its hidden instructions, producing content it should not, or behaving outside its role.

Indirect injection is usually more serious. The malicious instruction is planted somewhere the model will later read: a document uploaded for summarising, an email in an inbox the assistant processes, a web page it browses, or a record in the knowledge base it searches. The user may be entirely innocent; the attacker never talks to your system directly.

Why connecting AI to your data raises the stakes

A chatbot that can only chat has a limited blast radius. The risk grows with what the system can reach and do:

  • An assistant that searches your internal documents can be steered into surfacing content the current user should not see.
  • An assistant that reads incoming email can be instructed by the email itself.
  • An agent with tools (sending messages, updating records, calling APIs) can be instructed to use those tools on the attacker’s behalf.

This is why we always ask, early in any conversation, what a system can reach and what it is allowed to do. The same model can be low-risk in one deployment and high-risk in another.

What helps

There is no single fix, and no vendor can honestly promise immunity. What reduces the risk is layered design:

  1. Least privilege for agents. Give the system only the tools, data and permissions the task needs (OWASP calls the opposite excessive agency).
  2. Keep access control outside the model. Retrieval should only ever return what the current user is entitled to see, enforced by the system, not by asking the model to behave.
  3. Treat model output as untrusted input. Anything the model produces that reaches code, databases or browsers should be validated like any other user input.
  4. Put humans in the loop for consequential actions. Approval steps for payments, deletions or external messages.
  5. Monitor and test regularly. New techniques appear often; a system that was fine in spring may not be in autumn.

What testing can, and cannot, tell you

Independent testing shows how your specific system fails under realistic attack, at a point in time, within an agreed scope. A good report states the access the testers had, what they tried, what worked, how to fix it, and the risk that remains.

What testing cannot do is certify that a system is immune. Anyone who offers a “secure” stamp for an AI system is promising something the technology does not support. The useful questions are narrower and more honest: what did you test, how deep did you go, what did you find, and what is left?

If you are deploying an assistant or agent that can reach real data or take real actions, those are the questions to ask before it goes live.

Next step

Discuss your requirements.

Tell us what you want to build, or which AI system you need tested. We will tell you honestly whether we can do it well.

Talk to us

A scoping conversation first. Nothing starts until scope and limits are agreed in writing.