Service Line B · AI Security Assurance

Independent testing of the AI systems you rely on.

Red teaming, penetration testing and security assurance of AI systems, performed as a neutral third party. We only ever test systems we did not build.

What we test

From a single chatbot to a system of agents.

Testing scope grows with what the system can reach and do. Agents that use tools, and systems of agents, carry the most risk.

01

Chatbots and assistants

Injection, jailbreaks, leakage of hidden instructions and sensitive data.

02

Retrieval systems

AI that answers from company data: over-shared retrieval, cross-tenant access, poisoning.

03

Agents with tools

Excessive agency, insecure tool wiring, and actions taken on the wrong instructions.

04

Multi-agent systems

Trust between agents, chained failures, and paths from one agent’s compromise to real impact.

Every test also covers the classic surface around the model: application, API, identity and infrastructure. See the full method.

Depth of testing

More access buys deeper findings.

We test at the level of access you give us. Every report states that level, and what stayed out of reach.

Least access

Black box

We are givenNothing beyond what an outside attacker sees.

Finds bestExternal exposure, prompt injection from outside, jailbreaks.

Typical useA realistic attacker’s view; pre-launch checks.

The report statesCoverage was limited to external behaviour.

Partial access

Grey box

We are givenA normal user account and limited documentation.

Finds bestPrivilege abuse, leakage across roles, over-shared retrieval.

Typical useMost assistants and agents in production.

The report statesThe access given, and what stayed out of reach.

Full access

White box

We are givenSource, configuration, system prompts, data pipeline and architecture.

Finds bestDesign flaws, insecure tool wiring, poisoning paths, guardrail gaps.

Typical useHigh-risk or regulated systems.

The report statesThe full method; the highest confidence we offer.

How an engagement runs

Five steps. Nothing is touched until step two is signed.

Authorisation comes before access, always. It is our ethic and the legal footing for the work.

Gate requires a signature to proceed.

  1. 01
    ScopingWhat is in scope, the box level, and the access we will be given.
  2. 02
    Authorisation and rules of engagementSigned by someone with the authority to give it: timing, limits, stop conditions and AI-tool use, plus third-party permission where needed.
    GateNo access before this
  3. 03
    TestingAI-specific testing mapped to the OWASP Top 10 for LLM Applications and MITRE ATLAS, alongside application, API and infrastructure testing.
  4. 04
    ReportingFindings with severity, fix and residual risk; attack paths and a risk matrix; reasonable-not-absolute wording.
  5. 05
    RetestOptional verification that the fixes work, confirmed in a retest letter.

If we find something live, we stop.

If testing reveals a critical issue, live data or an active breach, we stop that test, preserve the evidence and tell your named contact at once. Testing resumes only when you approve it in writing.

Who it is for

For owners of AI systems, and the vendors who build them.

Organisations with an AI system

Built in-house, bought, or delivered by another vendor. You need to know how it fails before an attacker finds out.

AI build vendors

You want your work independently tested before delivery, so your client receives evidence, not assurances.

What you receive

Evidence you can act on, and limits you can see.

See a sample report
A findings report
What we tested, in what scope, when, what we found, how to fix it, and the residual risk.
An executive summary
One page: the headline risk, the risk matrix, and the three findings that matter most.
Attack paths
Chained findings drawn from entry point to impact, so the risk is seen, not described.
A retest letter (optional)
Verification that the fixes work, after you have applied them.

Our findings are point-in-time and reasonable, not absolute. We never call a system “secure”. Fixing findings is a separate engagement, and never one where we would then certify our own fix.

What we need from you

Signed authorisation and rules of engagement; a defined scope; and, for grey or white box testing, architecture documents, credentials or a test tenant. Where a cloud or platform provider’s terms require it, their written permission too.

What the rules of engagement cover

Ways to start

Start small, with a clear outcome.

Every engagement starts with a conversation about your needs and a scope agreed in writing. Each of these is a natural first step: narrow, concrete, and easy to approve.

A good first step

AI security assessment

One AI system, tested independently at the depth you choose, against the OWASP LLM Top 10 and the surface around it.

Time
Typically 1–2 weeks of testing for a single assistant; longer for agents and multi-agent systems.
You receive
Findings report with executive page, attack paths and residual risk.
Ask about this

After the first engagement

Ongoing assurance and support

Periodic re-testing as your system or its model changes, and maintenance for systems we built.

Time
Ongoing, on a schedule that matches how often your system changes.
You receive
Retest letters and an up-to-date view of residual risk.
Ask about this

FAQ

Questions we are often asked.

Something else on your mind? Ask us directly.

How long does an assessment take?

For a single assistant, typically one to two weeks of testing, plus time to write the report. Agents with tools and multi-agent systems take longer. We agree the window in the rules of engagement before we start.

Can you certify our system as secure?

No honest tester can. We report what we tested, in what scope, on what date, what we found, and the risk that remains. Our wording is point-in-time and reasonable, not absolute.

Can you independently test a system you built for us?

No. That is the one rule we never break. We can re-test our own builds as internal QA under a support agreement, but independent assurance on them goes to another firm.

What access do you need?

It depends on the depth you choose. Black box needs nothing beyond what an outside attacker sees; grey box needs a normal user account; white box needs source, configuration, system prompts and architecture. The report always states the level.

Will testing disrupt our production systems?

Techniques, test windows and exclusions are fixed in the rules of engagement before any access. Denial of service against production is excluded unless you explicitly agree otherwise in writing.

Do you use AI tools during testing? Where is our data processed?

Only with your written consent. The rules of engagement name the tools, what they are used for, where data is processed, and confirm that your data is not used to train them.

Do you fix the findings as well?

Fixing is a separate engagement. If we build a fix, we will not then give an independent verdict on it; a retest by us would be labelled internal QA.

Next step

Have an AI system that needs testing?

Tell us what it does, what it can reach, and when you need answers. We will propose a scope and a testing depth that fit.

Talk to us

A scoping conversation first. Nothing starts until scope and limits are agreed in writing.