Independent AI Assurance & Evaluation

Independent assurance for the AI making clinical decisions.

A diagnostic model that scores well in the lab can still fail at a new hospital, on an under-served patient group, or as the data shifts. We test diagnostic and predictive clinical AI for transferability, calibration, subgroup performance and drift, then map the evidence to the standards buyers and regulators trust.

Independent, audit-grade, reproducible, built for the UK regulatory landscape
What the AI Assurance Report delivers Evaluation scorecard Adversarial red-team findings Governance and compliance mapping Open Assurance Evidence Base
Why us

Trust in clinical AI shouldn't be a leap of faith.

The field of healthcare AI is long on principles and short on the assessment methods and oversight that turn principles into evidence. We are an independent assurance company built to close that gap: rigorous, reproducible testing that a clinical safety officer, a procurement lead or a regulator can actually follow.

Independent by design

We didn't build the model, so our evaluation has no reason to flatter it. That independence is exactly what makes the result credible to a buyer or a regulator.

Evidence over assurance theatre

Not slogans about responsible AI, but verifiable artefacts: an evaluation scorecard, red-team findings, and audit trails mapped to the frameworks that matter.

Reproducible and traceable

Every finding is reproducible and tied back to the evidence behind it, built on an open Assurance Evidence Base and tooling aligned with the UK AI Security Institute's Inspect.

What we test

Four tests a diagnostic model has to pass.

A practical evaluation framework for diagnostic and predictive clinical AI, anchored to the clinical reporting standards and UK clinical safety requirements (DCB0129 and DCB0160).

01

Transferability

Does performance hold at a new site, scanner or population, or does it quietly collapse outside the training data?

02

Calibration

When the model says 80 percent confident, is it right 80 percent of the time? Miscalibrated confidence misleads clinicians.

03

Subgroup performance

We break results down by age, sex, ethnicity and comorbidity, because a strong average can hide real harm to a specific group.

04

Drift over time

Models decay as practice and populations change. We test for drift and define the monitoring that keeps them safe in service.

See it in minutes

Why independent assurance matters

Listen instead

Independent assurance for clinical AI

The case for holding high-stakes clinical AI to a safety-critical standard, narrated for the commute, the queue, or between meetings. Generated with NotebookLM.

Questions, answered

What people ask about AI assurance

What is AI assurance for diagnostic AI?
AI assurance for diagnostic AI is the practice of independently verifying that a diagnostic or predictive model is safe, fair and reliable before and after deployment. It tests how the model transfers to new sites and populations, whether its confidence is well calibrated, how it performs across patient subgroups, and how it behaves as data drifts over time, then maps the evidence to the clinical reporting standards and governance frameworks that buyers and regulators rely on.
What does UK AI Evaluation deliver?
We deliver an end-to-end AI Assurance Report: an evaluation scorecard, adversarial red-team findings, and a mapping to the frameworks that matter, including TRIPOD+AI, STARD-AI, DECIDE-AI, CLAIM, DCB0129 and DCB0160. It is built on an open Assurance Evidence Base and tooling aligned with the UK AI Security Institute's Inspect.
Why focus on diagnostic and predictive clinical AI first?
Diagnostic and predictive models in radiology, pathology and haematology make high-stakes calls, yet many are validated only on the data they were trained on. Performance can fall sharply at a new hospital, on an under-represented subgroup, or as the population shifts. Independent transferability, calibration, subgroup and drift testing surfaces those failures before they reach a patient. Language-model red-teaming is a secondary track we also run.
What makes the assurance independent and audit-grade?
We are not the team that built the model, so our evaluation has no incentive to flatter it. Every result is reproducible and traceable to the evidence behind it, so a procurement lead, a clinical safety officer or a regulator can follow exactly how a conclusion was reached rather than take it on trust.
Who is UK AI Evaluation for?
We work with AI and healthtech vendors who need credible third-party assurance, with NHS trusts and procurement bodies assessing clinical AI before they buy, and with regulators and policymakers shaping the rules for safe medical AI in the United Kingdom.
Let's raise the bar together

Building, buying, or regulating clinical AI?

Whether you are a vendor who needs credible third-party assurance, an NHS trust assessing a model before you buy, or a regulator shaping the rules, let's talk about what independent assurance looks like for your context.