Black, grey or white box: choosing the depth of an AI security test
More access buys deeper findings. A practical guide to the three testing depths, what each one finds best, and how to choose.
When you commission a security test of an AI system, one of the first decisions is how much access to give the testers. The industry describes this as black, grey or white box testing. The choice shapes what the test can find, so it is worth understanding before you sign a scope.
Black box: the outside attacker’s view
The testers are given nothing beyond what any outside user would see: a public chat window, an API endpoint, an app.
Finds best: external exposure, prompt injection from outside, jailbreaks, and leakage of hidden instructions through normal use.
Good for: a realistic attacker’s-eye view, and quick pre-launch checks.
Limits: testers spend much of their time discovering how the system works rather than probing its weak points. Flaws deep in the design may never surface.
Grey box: a normal user, with some documentation
The testers get an ordinary user account, sometimes several with different roles, and limited documentation of how the system is meant to work.
Finds best: privilege abuse, leakage across roles and tenants, and over-shared retrieval, where an assistant returns documents the current user should not see.
Good for: most assistants and agents already in production. It is usually the best balance of depth and effort.
Limits: the testers still cannot see the system prompt, configuration or data pipeline, so some design-level issues are inferred rather than confirmed.
White box: full access
The testers see source code, configuration, system prompts, the data pipeline and the architecture.
Finds best: design flaws, insecure wiring between the model and its tools, data-poisoning paths into retrieval, and gaps in guardrails.
Good for: high-risk or regulated systems, and systems that act on the world through tools.
Limits: it needs more preparation and trust, and the findings depend on the documentation and access being complete.
How to choose
Three questions usually settle it:
- What can the system reach and do? The more data and tools it touches, the stronger the case for grey or white box.
- What decision rests on the result? A launch decision on a regulated system warrants more depth than a check on an internal prototype.
- What can you practically provide? White box testing needs people who can explain the design and grant access quickly.
What the report should say about depth
Whatever depth you choose, the report must say it plainly. A black box report should state that coverage was limited to external behaviour; a grey box report should state the access given and what stayed out of reach; a white box report should describe the full method.
Depth is not a technicality. It is part of what the findings mean. A clean black box result and a clean white box result are not the same statement, and a report that blurs the difference is not doing its job.