UK AI Evaluation
Policy engagement

Engaging with techUK's response to Ofgem's Call for Input on AI assurance in the energy sector

By Mohamed SG Omar, Founder of UK AI Evaluation · 3 August 2026

Ofgem has asked how AI in the energy system should be assured. techUK pulled together a member response, and we contributed. Here is what we put in, what landed, and the two recommendations we would still like to see.

What Ofgem asked, and what techUK did

Ofgem's Call for Input on AI assurance in the energy sector asks a practical question: as AI moves from proof of concept into live grid operations, forecasting, optimisation and control, how do we know it is operating safely, fairly and effectively? techUK drafted a collective response and asked members to comment on two rounds.

We responded to the first draft. Two sections had no answer text at all: Question 3 on critical infrastructure and the right level of rigour, and Question 7 on external assurance and standards. Both are squarely our subject, so we offered to help fill them and added three comments on the text as it stood.

What we contributed

The second draft took our points forward in substance. The parts we care about most:

Two recommendations we put to techUK

The response is much stronger, but two gaps remain if the goal is assurance a regulator can actually interrogate.

1. Make the evidence requirement concrete. For any material assurance claim, the record should identify the model and software version, the data or reference set, the test method, the assessment date, the operating assumptions, results by relevant subgroup, and the pre-agreed thresholds that trigger reassessment. That turns "continuous assurance" from a principle into something measurable: calibration drift, subgroup performance and distribution shift against a held-out reference set, with thresholds that force a re-evaluation.

2. Say what a credible assurance report must contain. An immature assurance market can produce cheap certificates that look like evidence and are not. Before treating any third-party verdict as assurance, guidance should set out the minimum: scope, limitations, independence and conflicts, methods, supporting evidence, findings, residual risks, validity period and re-evaluation triggers. A certificate or a headline verdict should not be mistaken for the underlying assurance.

Independent, reproducible assurance is what turns assurance theatre into something a regulator, a procurement lead or an operator can follow. The test is simple: can someone who did not build the system trace exactly how a conclusion was reached?

Why this matters

Energy is a high-consequence setting. An AI recommendation that is sound when generated can be wrong by the time it is executed, because operating conditions, network state, demand or cyber posture have moved. Assurance that only checks the model once, at launch, misses the part that actually fails. That is true in energy, and it is true in the diagnostic and predictive clinical AI we test every day: a model that performs well in the lab can quietly degrade at a new site, on an under-served subgroup, or as the data shifts.

The lesson travels across sectors. Assurance has to be continuous, versioned and independently checkable, not a certificate filed at deployment.

A note on independence

UK AI Evaluation is an independent AI assurance company. We are a techUK member and engaged with this Call for Input in that capacity. We are aligned with the direction of Ofgem, the UK AI Security Institute and the clinical safety frameworks we work to, but we are not affiliated with, certified by, or endorsed by any regulator or government body. Our interest is straightforward: better, checkable assurance standards benefit everyone who builds, buys or regulates AI in high-consequence settings.

Building or regulating high-consequence AI?

We test, evaluate and red-team diagnostic and predictive AI for transferability, calibration, subgroup performance and drift, then map the evidence to the standards buyers and regulators trust.