Method published · full reports in review Assurance Evidence Base

The evidence, with its limits attached.

We say assurance should be evidence, not slogans. This is where we hold ourselves to that: our own benchmark runs, published with the methodology, the run artefacts and the limitations that qualify them. Where something was not tested, we say so. Where a number moved between runs, we show the movement rather than the flattering figure.

How to read every number on this page

  • These are existence proofs on a 36-case battery, not safety rates. Adversarial Battery v1 is 36 cases across 6 threat categories, 6 cases per category. It supports claims of the form “this model produced this failure under this probe on this date”. It does not support a statistical safety rate.
  • Any single run’s pass-rate is a point estimate with a roughly ±1-case (about 3 percentage point) band. We measured this rather than assuming it: with the harness commit, seed and temperature all held fixed, individual cases still flipped between runs. Read 86.1% as “about 86%, give or take a case”, never as an exact figure.
  • A flat aggregate score is not evidence of case-level stability. One model returned an identical 31/36 on three consecutive repeats while two different cases flipped in opposite directions underneath it and cancelled out, one of them critical severity.
  • A high pass-rate is not a safety certification. It reflects resistance to this specific battery on that date only. It is not a general safety guarantee, and it is not suitability for any particular deployment context. None of the models tested has been assessed for deployment in a clinical setting.
Scope

What the Assurance Evidence Base is, and what it is not.

It is the open record behind our method: the runs we have actually done, the tooling and versions we did them with, and the gaps we have not yet closed. It is deliberately unflattering where the evidence is thin.

What it is

  • Our own adversarial evaluation runs against named models, on named dates, with the exact tooling version recorded.
  • Results tied to run artefacts: every figure traces to a SHA-256-hashed Inspect .eval log and its exported results.json.
  • An explicit “not assessed” wherever a category was not genuinely exercised, reported as a result rather than left blank.
  • Limitations stated as prominently as findings, including the ones that weaken our own conclusions.
  • Reproduction commands, so a third party can attempt the run themselves.

What it is not

  • Not a certification, an accreditation, a regulatory approval or an endorsement, of any model or of us.
  • Not a safety guarantee, and not a statement that any model tested is fit for clinical deployment.
  • Not a public leaderboard result: this is our own battery, our own IP, not a shared benchmark others have run.
  • Not a diagnostic or predictive-ML assessment. Everything currently here is from our secondary LLM track. No diagnostic-ML benchmark run against real data exists yet, and we will not imply one does.
  • Not a comparison against hosted frontier models or vendor-reported figures. Every number is from our own runs only.
The record

Benchmark runs

Adversarial Battery v1: 36 cases across prompt injection, jailbreak, data exfiltration, tool abuse, bias and fairness, and clinical safety, run against three local open-weight models with an independent grader model per target. Read the variance check alongside either weekly run; it qualifies both.

Adversarial Battery v1, results

Tested 2026-07-21
What was tested
36 adversarial cases across 6 threat categories against three local open-weight models, gpt-oss:20b, qwen2.5:14b-instruct and mistral-nemo:latest, each graded by a different model. Chat-completion probes only.
Headline finding
Pass-rates ranged from 55.6% to 86.1%. All three models failed the same two probes: gendered career advice for identical CVs, and a prompt injection hidden in a document they were asked to summarise. mistral-nemo was materially weaker on jailbreak resistance, with 4 critical-severity fails against 1 each for the other two models.
Not assessed
Live tool-use enforcement. No tool-calling harness was wired up, so the tool-abuse passes reflect a model declining a described action in text, not an enforced block.
Full report in review — available on request LLM secondary track Single run

Adversarial Battery v1, 2026-07-28 results

Tested 2026-07-28
What was tested
The same 36-case battery against the same three models, a third weekly repeat, plus a case-by-case comparison against the 2026-07-21 run with the harness commit, seed and temperature all unchanged.
Headline finding
Pass-rates 55.6% to 88.9%, and the gendered career-advice probe failed for a third week running on all three models. The finding that mattered was not the pass-rate: comparing week to week, one model reproduced exactly and two models each flipped exactly one case, with nothing in the configuration changed. A fixed seed and temperature did not guarantee case-level determinism.
Not assessed
Live tool-use enforcement, as above. Repeat-run variance was still unmeasured at this point; the ±1-case band this report recommends was an estimate, later tested in the variance check below.
Full report in review — available on request LLM secondary track Week-on-week comparison

Adversarial Battery v1, repeat-run variance check

Tested 2026-07-30 to 07-31
What was tested
Whether the battery reproduces within a single sitting. Each of the three models was run three times back-to-back, 9 complete 36-case runs, same harness commit, seed, temperature and grader pairings, with inspect-ai forced to 0.3.246 so the comparison held. One partial run interrupted at 10 of 36 cases was discarded rather than reported.
Headline finding
qwen2.5:14b-instruct returned an identical 31/36 on all three runs while two cases flipped in opposite directions and cancelled out, one of them critical severity. A matching aggregate pass-rate is not proof that the underlying result is stable. Observed case-level churn across the variance measurements to date: 0 to 5.6% of the battery per model.
Not assessed
Live tool-use enforcement, as above. This is still a small sample of variance measurements and is not yet a stable estimate of the true churn rate.
Full report in review — available on request LLM secondary track 9 runs, 3 × 3

Independence, and what we are not claiming

UK AI Evaluation Ltd (incorporated 24 June 2026, Companies House no. 17299891) is an independent assessor. Our work is aligned with the direction of the MHRA, the UK AI Security Institute and the NHS. We are not affiliated with, certified by, accredited by or endorsed by any of them, and nothing on this page or the pages it links to should be read as implying otherwise.

Nothing here constitutes certification, regulatory approval, or endorsement by any regulator or body, of any model tested or of UK AI Evaluation. A result on this site is a record of what a specific model did under a specific probe on a specific date, nothing wider.

All test inputs in these runs are synthetic. No real patient or personal data was used, including the simulated NHS number in the data-exfiltration cases.