Engaging with techUK's response to Ofgem's Call for Input on AI assurance in the energy sector
Ofgem has asked how AI in the energy system should be assured. techUK pulled together a member response, and we contributed. Here is what we put in, what landed, and the two recommendations we would still like to see.
What Ofgem asked, and what techUK did
Ofgem's Call for Input on AI assurance in the energy sector asks a practical question: as AI moves from proof of concept into live grid operations, forecasting, optimisation and control, how do we know it is operating safely, fairly and effectively? techUK drafted a collective response and asked members to comment on two rounds.
We responded to the first draft. Two sections had no answer text at all: Question 3 on critical infrastructure and the right level of rigour, and Question 7 on external assurance and standards. Both are squarely our subject, so we offered to help fill them and added three comments on the text as it stood.
What we contributed
The second draft took our points forward in substance. The parts we care about most:
- Questions 3 and 7 are now answered. The critical-infrastructure section scales rigour by operational authority, consequence and reversibility, and asks for bounded authority, execution-context validation, override, safe-state behaviour and reassessment after material change. The external-assurance section explains the value of independent evaluation and warns against false confidence from certification and one-off audits.
- The ISO/IEC 42001 distinction is clear. The draft states plainly that 42001 certifies the organisation, not a specific product or deployment. That single sentence does a lot of work, because a certificate alone tells a regulator nothing about the system that actually runs.
- Independent assurance is treated as different from self-attestation. The draft recognises that a verdict against a shared sector standard is more portable and trustworthy than organisations self-certifying against frameworks each interpreted differently.
Two recommendations we put to techUK
The response is much stronger, but two gaps remain if the goal is assurance a regulator can actually interrogate.
1. Make the evidence requirement concrete. For any material assurance claim, the record should identify the model and software version, the data or reference set, the test method, the assessment date, the operating assumptions, results by relevant subgroup, and the pre-agreed thresholds that trigger reassessment. That turns "continuous assurance" from a principle into something measurable: calibration drift, subgroup performance and distribution shift against a held-out reference set, with thresholds that force a re-evaluation.
2. Say what a credible assurance report must contain. An immature assurance market can produce cheap certificates that look like evidence and are not. Before treating any third-party verdict as assurance, guidance should set out the minimum: scope, limitations, independence and conflicts, methods, supporting evidence, findings, residual risks, validity period and re-evaluation triggers. A certificate or a headline verdict should not be mistaken for the underlying assurance.
Why this matters
Energy is a high-consequence setting. An AI recommendation that is sound when generated can be wrong by the time it is executed, because operating conditions, network state, demand or cyber posture have moved. Assurance that only checks the model once, at launch, misses the part that actually fails. That is true in energy, and it is true in the diagnostic and predictive clinical AI we test every day: a model that performs well in the lab can quietly degrade at a new site, on an under-served subgroup, or as the data shifts.
The lesson travels across sectors. Assurance has to be continuous, versioned and independently checkable, not a certificate filed at deployment.
A note on independence
UK AI Evaluation is an independent AI assurance company. We are a techUK member and engaged with this Call for Input in that capacity. We are aligned with the direction of Ofgem, the UK AI Security Institute and the clinical safety frameworks we work to, but we are not affiliated with, certified by, or endorsed by any regulator or government body. Our interest is straightforward: better, checkable assurance standards benefit everyone who builds, buys or regulates AI in high-consequence settings.
Building or regulating high-consequence AI?
We test, evaluate and red-team diagnostic and predictive AI for transferability, calibration, subgroup performance and drift, then map the evidence to the standards buyers and regulators trust.