We think it should be easier to reliably answer the basic questions: what AI systems are public bodies using? How are those systems affecting people?
To help civil society actors work out how to get better answers to these questions, we brought together people working across academia, law, journalism, technology and public policy to discuss the challenges involved in scrutinising government use of AI.
We asked about what useful infrastructure already exists, and where the gaps are, and where participants identified shared needs across civil society.
The resulting discussion revealed a series of useful tensions: between collecting more information and making better use of what already exists; between understanding how a system works and examining what it actually does; and between involving affected communities and ensuring that governments remain responsible for proper oversight.
Do we lack information, or is it simply difficult to find?
In many cases we still do not have enough basic information about government AI use. Existing transparency mechanisms do not give us a comprehensive picture of which systems are being developed, trialled or used across government.
The UK government’s Algorithmic Transparency Recording Standard is widely seen as ‘good practice’, and could provide valuable information if used consistently, but we know that not every relevant system appears on the public hub.
Even then, the healthiest possible central register is unlikely to capture every way AI is used in practice. Alongside formally procured systems, staff may use generative AI for drafting, research, transcription or coding. Public bodies themselves may not have a complete overview of these kinds of informal AI use.
From some participants we heard that the problem isn’t actually lack of data, it’s what to do with the data. There is already considerable information available if you know where to look. Details emerge through procurement notices, court cases, policy documents, impact assessments, transparency publications and Freedom of Information requests. The problem is that this information is fragmented, and that it isn’t clear how to use it to influence the decisions made.
These are not necessarily opposing accounts. They point to several different failures:
- information that public bodies do not collect
- information that is collected but not published
- information held by private suppliers rather than public authorities
- information that is public but fragmented or difficult to find
- and information that is published without enough context to support scrutiny
- a lack of mechanisms for collating and effectively using information on impact
Saying that we “need more data” risks collapsing all of these problems into one. Before deciding on a solution, we need to understand where in the chain the information is being lost.
What counts as AI, and does the definition help us?
We debated how important it is to define what we mean by AI. There’s a risk that by focusing too much on what is or isn’t AI, we offer a get-out clause. A public body may argue that a system is not really AI, or has a human checking its output, therefore there is no cause for concern.
The example of an AI transcription tool used in social care was raised. On paper, transcription might look like a low-risk administrative task. In practice, the resulting record could influence how a conversation is understood and what decisions are subsequently made about a child or family. These transcribed accounts could become really important later down the line. Similarly, an AI tutoring tool may not make formal decisions about pupils, but its widespread use could still have major effects on children’s education.
Rather than waiting for a perfect definition of AI, it may be more useful to ask what the technology is doing, how people use its output, who is affected and what happens when something goes wrong.
How should we evaluate these systems?
In the room there was a lot of agreement around the need to focus on evidence about how the systems work in reality, how they affect real citizens, and especially the most marginalised.
However ‘evaluation’ can mean different things to different people. The existence of an evaluation does not necessarily mean that meaningful scrutiny has taken place. For example, technology suppliers may assess their own products using benchmarks that make sense to developers but not to the wider public or civil society. A claim that a system is “98% accurate” is difficult to interpret without knowing what was tested, which mistakes make up the remaining 2%, who experiences those mistakes and what the consequences are.
There is a risk of “evaluation washing”: the language of evaluation lending legitimacy to a system without answering the questions that matter.
What kinds of context help make an evaluation meaningful? Perhaps it should consider not only technical performance, but how a tool interacts with staff, existing services and the people affected by it. A decision-support tool may perform well in a controlled test but operate very differently when used by overstretched staff who have little time or authority to question its recommendations.
As part of this conversation about a focus on what systems actually do, there was some disagreement about how much attention should be paid to explaining how systems work. Focusing on impact helps rebalance the conversation towards what actually happens to people, but technical information can be necessary for researchers, lawyers and affected individuals in some cases. It may help explain why errors occur, whether particular groups are treated differently and where responsibility lies for this.
Connecting official records with frontline experience
Official records can tell us what public bodies say they are doing, but this is really brought to light by frontline organisations who can help reveal what happens in practice.
Caseworkers are often among the first to see the effects of new technology. They may encounter individual cases that appear isolated but form part of a larger pattern. We heard a powerful example from the domestic violence sector when 20 charities each reported one instance of a new trend: this helped both reveal the recurring issue and suggest a possible solution that no organisation could identify alone.
However, organisations do not always know what to look for or how to establish whether AI has contributed to a new problem. Changes may first appear as an unexplained decision, a new administrative barrier, or a small but specific reduction in service quality. Better connections between frontline organisations and more technically minded organisations could help turn these scattered experiences into evidence of systemic problems.
We heard some debate about how much to involve affected communities. Whilst real people should be at the heart of our thinking, participation must not transfer too much responsibility onto communities, especially those already vulnerable. People should not need to become experts in every system affecting them in order to be protected. Public bodies remain responsible for demonstrating that the technologies they use are appropriate, safe, fair and effective. Public participation should strengthen oversight, not replace it.
How do we move forward?
What would we like to see?
- We need to know what systems public bodies are developing and using, and for this to be reported as structured data. Reporting hubs for this purpose exist but are not well-used.
- We need information early enough for people to influence decisions, not after a system has been introduced. This requires better transparency throughout procurement and design processes.
- We need evaluations that examine real-world outcomes rather than relying on abstract benchmarks. Civil society has a role to play here in shaping these evaluations.
- We need practical ways to use that evidence to challenge individual decisions and wider patterns of harm. This requires sector-wide collaboration to share knowledge, cases, challenges and successes.
The task is not simply to publish more data (although that is a key request to public bodies). It is to connect official records with real-world experiences and create meaningful routes to challenge and improve. Transparency matters because of what it allows people to do.
Huge thanks to everyone who participated in this roundtable discussion. We’ll be using the outputs to help shape our future research and digital service development in this area.
This roundtable formed part of work supported by the Joseph Rowntree Charitable Trust.
—
Image: Jamillah Knowles via Better Images of AI (CC-BY 4.0)