Approach · How we test

A method you can inspect before you hire us.

What we test, how deep we go, the rules every engagement is signed under, and the standard every report meets.

Coverage

All ten 2026 OWASP risks for LLM applications, plus the surface around the model.

We plan adversary techniques against MITRE ATLAS and test the application, API, identity and infrastructure the model sits within.

  1. LLM01Prompt InjectionDirect and indirect injection, including through documents, web pages and tool outputs.
  2. LLM02Sensitive Information DisclosureSecrets, personal data and confidential content leaking through outputs.
  3. LLM03Excessive AgencyAgents with more tools, permissions or autonomy than the task needs.
  4. LLM04Supply ChainThird-party models, plugins, datasets and dependencies.
  5. LLM05Data and Model PoisoningTampered training, fine-tuning or retrieval data.
  6. LLM06Unbounded ConsumptionCost and denial of service through resource abuse.
  7. LLM07MisinformationConfident false output in decisions that matter.
  8. LLM08Hidden Context ExposureExtraction of system prompts and anything placed in front of the model unseen.
  9. LLM09Vector and Embedding WeaknessesRetrieval access control, cross-tenant retrieval, embedding inversion.
  10. LLM10Improper Output HandlingModel output reaching code, queries or browsers unsanitised.
  11. +Beyond the modelApplication, API, identity and infrastructure: the classic attack surface around the model, with adversary techniques planned against MITRE ATLAS.

Source: OWASP GenAI Security Project, Top 10 for LLM Applications 2026.

An engagement, step by step

From first call to retest letter.

  1. 01Scoping
  2. 02Authorisation
  3. 03Testing
  4. 04Reporting
  5. 05Retest
  1. Step 1

    Scoping

    We agree what is in scope, the depth (black, grey or white box) and the access you will give us. Out-of-scope systems are named explicitly.

    You receiveScope statement

  2. Step 2

    Authorisation

    Someone with the authority to approve testing signs the rules of engagement: timing, limits, stop conditions and AI-tool use. Third-party permission where needed.

    You receiveSigned rules of engagement · no access before this

  3. Step 3

    Testing

    AI-specific testing mapped to the OWASP LLM Top 10 and MITRE ATLAS, alongside application, API and infrastructure testing. If we find something live, we stop and tell you.

    You receiveEvidence log

  4. Step 4

    Reporting

    Findings with severity, evidence and fixes; attack paths; a risk matrix; and the residual risk, in point-in-time, reasonable-not-absolute wording.

    You receiveFindings report

  5. Step 5

    Retest

    Once you have applied fixes, we verify that they work and confirm it in writing. Optional, and recommended for anything rated high or critical.

    You receiveRetest letter

Depth

Black, grey or white box, and the report always says which.

Least access

Black box

We are givenNothing beyond what an outside attacker sees.

Finds bestExternal exposure, prompt injection from outside, jailbreaks.

Typical useA realistic attacker’s view; pre-launch checks.

The report statesCoverage was limited to external behaviour.

Partial access

Grey box

We are givenA normal user account and limited documentation.

Finds bestPrivilege abuse, leakage across roles, over-shared retrieval.

Typical useMost assistants and agents in production.

The report statesThe access given, and what stayed out of reach.

Full access

White box

We are givenSource, configuration, system prompts, data pipeline and architecture.

Finds bestDesign flaws, insecure tool wiring, poisoning paths, guardrail gaps.

Typical useHigh-risk or regulated systems.

The report statesThe full method; the highest confidence we offer.

Rules of engagement

Ten sections, signed by every counterparty before any access.

The rules of engagement are the authority we show: to our own people, to your cloud and platform providers, and to the AI tools we test with. They describe exactly what you agreed to, and nothing more.

Sending your data or screenshots into an AI tool is data processing. Section 05 records your written consent, in line with the UAE PDPL.

  1. 01
    Parties and authorityWho signs, their authority to authorise testing, and the named testers.
  2. 02
    Scope and exclusionsIn-scope systems, AI components and environments; what is explicitly off limits.
  3. 03
    Test windowDates, hours, blackout periods and time zones.
  4. 04
    TechniquesWhat is permitted and what is prohibited: no denial of service on production, no social engineering unless named.
  5. 05
    AI-assisted testingWhich AI tools testers may use and for what; your consent to process test data in them; no training on your data; where the data is processed.
  6. 06
    Data handlingHow evidence, credentials and findings are stored, redacted, retained and destroyed.
  7. 07
    Third partiesWritten permission from cloud and platform providers where their terms require it.
  8. 08
    Communication and stopContacts, escalation, stop conditions and what happens if we find something live.
  9. 09
    ReportingInterim and final reports, their format and who receives them.
  10. 10
    Sign-offSigned by every counterparty before any access. No signature, no test.

Stop conditions

If we find something live, testing stops and you decide what happens next.

The one call a tester makes alone is to stop. Every decision after that belongs to you.

  1. TriggerCritical issue, live data, or an active breachFound during authorised testing.
  2. Step 1Stop that testPreserve the evidence; change nothing further.
  3. Step 2Tell your named contactAt once, by the channel set in the rules of engagement.
  4. Step 3Resume only with written approvalYou decide; we record the decision.

Anything below that threshold is logged, scored with evidence and reported as normal.

How we report

Our reports are a product in their own right.

Every assurance report follows the same standard, so you can compare one engagement with the next.

See a sample report
  • One page for executives. The headline risk, the risk matrix and the three findings that matter most.
  • Attack paths. Chained findings drawn from entry point to impact, so the risk is seen, not described.
  • Evidence that stands up. Annotated screenshots with secrets and personal data redacted, and steps to reproduce.
  • Built from data. Charts and diagrams generated from the findings themselves, never screenshots of charts.
  • Honest wording. Point-in-time, reasonable not absolute, with the residual risk stated. Severity is always given with the scope and date of the test.

Likelihood ↑Impact →

Critical 15–25 · Fix before go-live, or immediately if live.High 10–14 · Fix within days.Medium 5–9 · Fix in the next release.Low 1–4 · Fix as capacity allows.

Next step

Discuss your requirements.

Tell us what you want to build, or which AI system you need tested. We will tell you honestly whether we can do it well.

Talk to us

A scoping conversation first. Nothing starts until scope and limits are agreed in writing.