-
We think it should be easier to reliably answer the basic questions: what AI systems are public bodies using? How are those systems affecting people?
To help civil society actors work out how to get better answers to these questions, we brought together people working across academia, law, journalism, technology and public policy to discuss the challenges involved in scrutinising government use of AI.
We asked about what useful infrastructure already exists, and where the gaps are, and where participants identified shared needs across civil society.
The resulting discussion revealed a series of useful tensions: between collecting more information and making better use of what already exists; between understanding how a system works and examining what it actually does; and between involving affected communities and ensuring that governments remain responsible for proper oversight.
Do we lack information, or is it simply difficult to find?
In many cases we still do not have enough basic information about government AI use. Existing transparency mechanisms do not give us a comprehensive picture of which systems are being developed, trialled or used across government.
The UK government’s Algorithmic Transparency Recording Standard is widely seen as ‘good practice’, and could provide valuable information if used consistently, but we know that not every relevant system appears on the public hub.
Even then, the healthiest possible central register is unlikely to capture every way AI is used in practice. Alongside formally procured systems, staff may use generative AI for drafting, research, transcription or coding. Public bodies themselves may not have a complete overview of these kinds of informal AI use.
From some participants we heard that the problem isn’t actually lack of data, it’s what to do with the data. There is already considerable information available if you know where to look. Details emerge through procurement notices, court cases, policy documents, impact assessments, transparency publications and Freedom of Information requests. The problem is that this information is fragmented, and that it isn’t clear how to use it to influence the decisions made.
These are not necessarily opposing accounts. They point to several different failures:
- information that public bodies do not collect
- information that is collected but not published
- information held by private suppliers rather than public authorities
- information that is public but fragmented or difficult to find
- and information that is published without enough context to support scrutiny
- a lack of mechanisms for collating and effectively using information on impact
Saying that we “need more data” risks collapsing all of these problems into one. Before deciding on a solution, we need to understand where in the chain the information is being lost.
What counts as AI, and does the definition help us?
We debated how important it is to define what we mean by AI. There’s a risk that by focusing too much on what is or isn’t AI, we offer a get-out clause. A public body may argue that a system is not really AI, or has a human checking its output, therefore there is no cause for concern.
The example of an AI transcription tool used in social care was raised. On paper, transcription might look like a low-risk administrative task. In practice, the resulting record could influence how a conversation is understood and what decisions are subsequently made about a child or family. These transcribed accounts could become really important later down the line. Similarly, an AI tutoring tool may not make formal decisions about pupils, but its widespread use could still have major effects on children’s education.
Rather than waiting for a perfect definition of AI, it may be more useful to ask what the technology is doing, how people use its output, who is affected and what happens when something goes wrong.
How should we evaluate these systems?
In the room there was a lot of agreement around the need to focus on evidence about how the systems work in reality, how they affect real citizens, and especially the most marginalised.
However ‘evaluation’ can mean different things to different people. The existence of an evaluation does not necessarily mean that meaningful scrutiny has taken place. For example, technology suppliers may assess their own products using benchmarks that make sense to developers but not to the wider public or civil society. A claim that a system is “98% accurate” is difficult to interpret without knowing what was tested, which mistakes make up the remaining 2%, who experiences those mistakes and what the consequences are.
There is a risk of “evaluation washing”: the language of evaluation lending legitimacy to a system without answering the questions that matter.
What kinds of context help make an evaluation meaningful? Perhaps it should consider not only technical performance, but how a tool interacts with staff, existing services and the people affected by it. A decision-support tool may perform well in a controlled test but operate very differently when used by overstretched staff who have little time or authority to question its recommendations.
As part of this conversation about a focus on what systems actually do, there was some disagreement about how much attention should be paid to explaining how systems work. Focusing on impact helps rebalance the conversation towards what actually happens to people, but technical information can be necessary for researchers, lawyers and affected individuals in some cases. It may help explain why errors occur, whether particular groups are treated differently and where responsibility lies for this.
Connecting official records with frontline experience
Official records can tell us what public bodies say they are doing, but this is really brought to light by frontline organisations who can help reveal what happens in practice.
Caseworkers are often among the first to see the effects of new technology. They may encounter individual cases that appear isolated but form part of a larger pattern. We heard a powerful example from the domestic violence sector when 20 charities each reported one instance of a new trend: this helped both reveal the recurring issue and suggest a possible solution that no organisation could identify alone.
However, organisations do not always know what to look for or how to establish whether AI has contributed to a new problem. Changes may first appear as an unexplained decision, a new administrative barrier, or a small but specific reduction in service quality. Better connections between frontline organisations and more technically minded organisations could help turn these scattered experiences into evidence of systemic problems.
We heard some debate about how much to involve affected communities. Whilst real people should be at the heart of our thinking, participation must not transfer too much responsibility onto communities, especially those already vulnerable. People should not need to become experts in every system affecting them in order to be protected. Public bodies remain responsible for demonstrating that the technologies they use are appropriate, safe, fair and effective. Public participation should strengthen oversight, not replace it.
How do we move forward?
What would we like to see?
- We need to know what systems public bodies are developing and using, and for this to be reported as structured data. Reporting hubs for this purpose exist but are not well-used.
- We need information early enough for people to influence decisions, not after a system has been introduced. This requires better transparency throughout procurement and design processes.
- We need evaluations that examine real-world outcomes rather than relying on abstract benchmarks. Civil society has a role to play here in shaping these evaluations.
- We need practical ways to use that evidence to challenge individual decisions and wider patterns of harm. This requires sector-wide collaboration to share knowledge, cases, challenges and successes.
The task is not simply to publish more data (although that is a key request to public bodies). It is to connect official records with real-world experiences and create meaningful routes to challenge and improve. Transparency matters because of what it allows people to do.
Huge thanks to everyone who participated in this roundtable discussion. We’ll be using the outputs to help shape our future research and digital service development in this area.
This roundtable formed part of work supported by the Joseph Rowntree Charitable Trust.
—
Image: Jamillah Knowles via Better Images of AI (CC-BY 4.0)
-
Authors: Ben Worthy (Birkbeck), Laszlo Horvath (Birkbeck), Julia Cushion (mySociety), Alex Parsons(mySociety)
—
Artificial Intelligence, as we know, is everywhere. For governments and public bodies, the possibilities for policy-making seem endless. As the UK government put it in its AI playbook ‘the potential of AI to transform public services is enormous, giving us an unparalleled opportunity to do things differently and deliver more with less’.
As with every other part of government, AI use has spread across local government in the UK. According to the Local Government Association around 70% of councils are now using them in frontline services, with almost all councils ‘using or exploring AI’ in some way.
What is less clear is what AI is exactly being used for, and how far advanced the work is. Are these experiments? Will pilots be abandoned? More importantly, how easy is it to make AI assisted policies democratically accountable? It is not clear how transparent and accountable these new AI policies are, or how they can be scrutinised.
Looking at London
London’s 32 boroughs are at the forefront of AI use and adoption. This 2025 survey put London’s local government at the top of the UK innovators, alongside Scotland. The London Office of Technology and Innovation (LOTI) is co-ordinating work and ideas, and several councils have partnerships with universities.
Our joint project tried a first mapping of London’s local government to try and see more clearly where AI tools are being used in frontline policy areas and to get a sense of how transparent they are. To map the landscape, we used a combination of methods, including use of primary published official documents (strategy papers, charters), media searches and surveys (such as by the LGA), as well as logs of FOI requests already made. We hope this gives a first insight into where AI policy is happening, and how open and accountable it is to the public.
Where is AI happening?
In line with the wider picture, the most popular area for AI use in London is around Adult Social Care, particularly with use of AI for notes and minute takers. This isn’t surprising given that, for every £1 local councils spend, 39p of it goes on adult social care. The second most popular area for AI use was housing, then followed by AI chatbots for citizen contact. There was then a long tail of rather varied use from CCTV to planning.
It needs to be remembered that not all of these innovations are equal in size: for example Hillingdon’s AI chatbot covers 40% of all citizen interaction. Measuring the exact impact of each policy is complex, as some supposed ‘backroom’ services may make a difference to front-facing policy, as this experiment showed.
Areas of AI use across London Boroughs
Area Sub Area Occurrence Adult Social care Notes, pain checker, alarms 12 Housing Mould, HMO, Rent, Complaints 7 AI Chatbot citizen contact 5 Backroom 4 AI Chatbot translation 3 AI Children services 2 Planning Analysis of applications 2 Surveillance CCTV 2 Traffic Potholes, Monitoring 2 Waste Fly tipping 2 FOI Chatbot, collate 2 So, at least for London, we can see that AI policy is happening, mainly across particular policy areas and in certain forms, but with experiments and innovations spreading across new areas.
How transparent and accountable are AI policies?
We then tried to map and trace various ways in which you can potentially access information about these policies. Some may be proactively published by councils — such as policy documents— and some are legal routes to access— such as FOI or Data protection. Others are voluntary/mandatory schemes— like the Algorithmic Transparency Recording Standard (ATRS) (not mandatory for local government but advised).
In 2024 the National Audit Office were concerned that accountability was a bit of a ‘patchwork’. Our analysis found the same, that there is a patchwork of transparency approaches. There are some signs of direct access, with published data (via data strategies) and FOI being the most common ways to find out. About a third of authorities have a published strategy/charter; eight have received FOI requests about this, and eight had mentions of it in meetings. The local media have played a role with stories. In terms of indirect access other less focused forms included publicity, trade media reporting and internal oversight from debates and council meetings. These tools created in turn varying degrees of data and transparency.
Transparency tools for AI use across London boroughs
AI Use Council count Notes Oversight Board 4 3 others ‘in progress’ Joint Working Bodies (with University) 2 Transparency documents (documents/strategy) 11 Recorded on Algorithmic Transparency Recording Standard (ATRS
4 [Plus GLC} Transparency (stories in local media) 16 Transparency (FOI Log with FOI requests on AI) 8 61 FOI requests in total Transparency (mention in council meetings, committees etc) 8 Despite the London Office of Technology recommending a specific board of experts, few authorities have them (yet), with just four having one in place. The overall picture again is patchy and unclear – there’s no one fixed mechanism being put in place or used.
The patchy landscape of AI transparency in London is somewhat troubling if we consider that, in many respects, London is the least demanding test for public sector transparency. Its boroughs have been experimenting with algorithmic tools for some years, and that experimentation has unfolded alongside comparatively well-developed local scrutiny, for example through local media or via London’s universities. If a minimum set of transparency practices were going to take hold anywhere, one would expect it to be here with reasonable institutional capacity.
—
Image: David Monaghan
-
The effects of AI are making themselves known across every sector, and Freedom of Information is no exception.
AI has lowered the bar to making contact, in every place where citizens interact with public authorities, from making pothole reports to submitting FOI requests.
On the face of it, this is a good thing for individuals — after all, mySociety was founded on the principle that digital processes should make it easier for people to interact with government — but this sudden and disruptive uptick in communication has created an increased burden for those authorities*.
At FOI Fest, Scottish FOI Commissioner David Hamilton told us that they’re experiencing unprecedented levels of requests, due to a number of factors, but certainly in part because people are using AI — and as he explained, these requests bring their own specific types of challenges.
Now the UK’s Information Commissioner’s Office, the ICO, has addressed this new reality by including advice around using AI in their wider guidelines for making an effective request (which also, we’re glad to see, point people towards WhatDoTheyKnow).
What’s the problem?
Generative AI tends to add more complex and unnecessary wording, can provide inaccurate information and will sometimes ‘hallucinate’ (provide inaccurate information). All of these factors mean that authorities are having to spend more time corresponding with request-makers to pin down what they meant to ask for.
While it need not be, FOI can already sometimes be a long process, and this only adds to the time taken (when an authority needs to ask for clarification, it restarts the clock on the 20 working days within which they are required to provide a response).
Plus, if you end up needing to make a complaint about the way your request has been handled, the ICO notes that complex and inaccurate requests make it much harder for them to help you.
What you need to know
The commissioners’ offices are not saying that AI should never be used to write an FOI request or appeal; simply that if you do so, you should proceed with due caution. They suggest:
- Checking that AI hasn’t changed the scope of what you have asked for, and that you are asking for precisely — and only — the information you require.
- Ensuring that your request is clear, concise and focused.
- Checking the facts: don’t assume AI “knows” which law applies or which authority is the suitable recipient for your request.
- Ensuring the tone is appropriate.
In many ways, this is the same advice we’d apply for request-makers who are using their own brains, as well. Either way, it’s always worth getting it right, and making things as easy for yourself, and our public authorities, as possible.
*Although, NB that the jury is still out on whether AI is the direct cause of this: see, for example, this story from Australia.
—
Image [amended to include ‘Make an FOI request’ button]: Zulfugar Karimov
-
LLMs can increase demand on public systems by removing the friction that previously limited access. One potential result of this is new forms of unconsidered rationing that recreate that friction. Instead, we should move away from zero sum systems and aim for technical and policy approaches that turn unscalable private benefits into efficient collective ones.
Many kinds of citizen-driven interactions with the public sector are rationed through friction: fewer people engage in them than might do otherwise, because they feel the process is time-consuming or requires expertise. We can see examples of this in planning objections, correspondence with elected representatives, FOI requests and consultation responses. LLM technologies can lower the time or expertise required and also prompt people to engage in the processes in the first place: “Would you like me to draft a complaint about this?”
Systematic impacts
This reduced friction may be good for individuals, but the resulting increase in engagement can overwhelm the system itself. In response, it may slow down, collapse, or adopt new means of rationing or prioritising access. We can see indications of this across different kinds of interactions: journalist Martin Rosenbaum has identified an upward trend across public sector complaints organisations, and concerns are being raised across sectors about AI’s contribution to growth in volumes.
So how should organisations that handle public submissions respond in an informed way? Here’s an approach to thinking about the problem. We can divide these interactions into three types:
- Private benefit – when an interaction has a benefit almost exclusively to the requester, either competitively (eg a grant or job application), or non-competitively (eg an application for a state benefit).
- Collective benefit – when an interaction has a benefit to the requester, and also to wider society (eg a public FOI request, reporting a pothole).
- Zero sum interaction – when an interaction success for one person is a failure for another (eg planning).
Private benefits
Some public services fall clearly in the first category: they are unavoidably a collection of private interactions. For these, there might be improved efficiencies to be found in delivery at scale but, particularly for non-competitive benefits, these are also likely to eventually run into decisions either about increasing provision (assuming a higher level of claims from those entitled going forward), or new forms of rationing.
As stands, AI inputs can both improve the efficiency of systems through sharper, more complete initial submissions, but can also make more verbose and complex submissions that cite non-existent law. To prioritise the former over the latter, systems can explore triage approaches that enforce or encourage the qualities that make input valuable: clarity, accuracy and concision.
When running into real limits, it is important to be clear about the criteria you want to ration on, and that they are in line with the overall purpose of the system, rather than implicitly prioritising those with greater resources. In their FOI complaints system, the ICO is using public benefit as a criteria for prioritisation. The British Academy uses partial randomisation above a scoring cutoff to ration randomly rather than requiring additional work (on both sides) to further differentiate.
Collective benefits
A bigger win is, where possible, to transform private benefits into collective benefits. In these cases, reduced friction is self-regulating because spillover benefits from an individual’s case help reduce demand from others : the private benefit person B is looking for has already been provided by person A’s interaction.
One of the key ways mySociety’s services help people is to harness the self-interest of individual users for collective benefits. Every public request made on WhatDoTheyKnow also adds to the pool of public knowledge accessible on the internet, reducing the need for duplicate requests (with a similar logic to reducing duplicate reports on FixMyStreet). This means we can effectively lower the bar to access while improving overall efficiency of the system.
We come to this from a technology lens, but the same principles apply from an institutional-design approach. For instance, if MPs’ casework or complaints are increasing, you want to shift towards more systematic rather than individual benefits from casework. This looks like support for better collective learning, and an improved ombudsman to support collective rather than individual fixes. This kind of approach works best where good statistics are collected at a system level to help identify what collective changes are needed: tracking the overall level of demand, level of demand to different parts of the system and nature of the demand, ie what are people asking for.
Zero sum systems
The biggest shift needed is in reforming zero-sum systems, where there is currently an incentive for both sides to escalate the volume. Reduced friction here just raises costs for all concerned rather than giving increased benefits to anyone. Individual use of AI to create submissions is individually enabling in these cases, but not collectively. So, in the words of the 1980s classic film War Games, “the only winning move is not to play”. The real innovation is in solutions that open up new, and more effective, ways of working out what everyone can live with, rather than recreating rationing through new means. For instance, rather than adversarial AI planning objection generators, we could aim for a collaborative planning system that through improved communication and coordination lowers costs and removes incentives to volumes of engagement.
Red flags for zero sum interactions are when volume is implicitly being used as a proxy for strength of feeling, or popularity of a particular viewpoint, because its value as a signal is going to become increasingly degraded as AI use increases.
Systems work better when the benefits are collective rather than atomised
Mass adoption of AI removes one set of bottlenecks, but this can create capacity challenges for public systems. Previous waves of civic technology have built on reduced costs of storing and sharing information to build systems that help share the benefits of people’s work and lower the barriers to entry.
The current wave of AI chatbots cut against this, encouraging atomised approaches, rather than collective ones. We need to explore technical and policy approaches that help systems better achieve their purpose, without giving up on the idea of lowering barriers to entry. We can do this both by exploring how the technological features of AI tools can be bent towards collective gains, and moving away from systems that incentivise these approaches.
—
Image: Engin Akyurt
-
Generative AI is good at solving some kinds of problems, and bad at solving others. With the rush to apply AI approaches across the public and private sector, we want to encourage people to use the right tool for the right problem. This blog post proposes a test that makes it easy to understand whether or not the applications are genuinely beneficial for the job in hand.
Generative AI has no concept of truth. It is designed to create outputs that are internally consistent, and this might or might not coincide with true things when the training data and context are well aligned. By now, we’ve all heard examples of false-positive hallucinations, where AI has asserted that something exists or was said because doing so is internally consistent with the question — but which turns out not to be true. Depending on the application, if unchecked, this can have catastrophic effects, meaning that validation of outputs is essential.
How to assess your project for AI suitability
In our recent Shifting Landscapes report, we shared a simple matrix that helps to assess how useful it is to apply an AI approach to any given problem.
It asks how hard/expensive is it currently to produce a solution without AI, and how hard/expensive is to verify that the solution is correct, with four potential outcomes:
Producing a solution is cheap/easy Producing a solution is hard/expensive Verifying the solution is cheap/easy Weak AI benefits (which may increase at scale) Significant AI benefits Verifying the solution is hard/expensive Get a human to do it Break down the verification problem (and repeat) Let’s look at each possible outcome in turn:
1. Weak AI benefits (which may increase at scale)
producing a solution is cheap / verifying the solution is cheapThis applies to tasks where AI tools might help people complete tasks more efficiently, but where the resulting impact or time savings are not significant. Over time/mass use, the benefits might increase.
Examples here include tasks like letter-writing and making summaries of documents or transcripts. If AI can do the initial grunt work, a human can take over and make tweaks to the output, nominally saving some time.
In our own field of civic tech, we can see this kind of tool being used to help people navigate bureaucracy: it might help format letters to representatives, or make effective appeals when FOI requests are refused.
Cheap processes at scale can also unlock new collective benefits. For instance, Muckrock uses LLMs to extract information and success/fail status from individual FOI responses. Doing this manually per request is easy for people, but requires lots of people to do the work to create a useful dataset across the entire corpus. An AI approach drops the costs further, which produces a small benefit on an individual scale, but collectively creates useful data.
As we note in our AI Framework, we have to recognise that a large number of small uses can build up into a negative effect. For instance, AI-created objections to planning applications might overwhelm a system that was built for a world in which there are higher hurdles to lodging an objection.
2. Significant AI benefits
producing a solution is expensive / verifying the solution is cheapIn this scenario, we’re thinking of situations where it is harder for a human to create a credible solution than it is to check if the outputs are valid. Conceiving a solution might be hard because it requires specialised knowledge, such as coding, or significant time and resources, like the analysis of a huge dataset; but it would be easy for a human to see whether or not the solution is working as intended.
One of the biggest practical uses of AI so far has been seen in coding, because coding problems fit so well into this category, and so provide potential benefits. The structure of computer code is often formally checkable (for at least syntax errors), and often there is a relatively short turnaround between “having code” and “checking the code is effective”. This isn’t to say that all coding fits in this box, but enough that a clearly productive set of tools exists.
There are strong potential benefits here because an expensive process can be made cheaper, while the quality of the output can be checked through relatively cheap verification methods.
This segment of applications can be impactful even where access to models is relatively expensive, as a relatively small number of LLM users can have a big impact through the products that emerge.
3. Get a human to do it
producing a solution is cheap / verifying the solution is expensiveSome LLM processes produce outputs that cannot be quickly verified by automatic or human means.
Here, using an LLM for the initial solution might be less effective than having a human do it from the start. While tweaking an email that contains slightly poor wording is a cheap correction, adjusting a multi-page report written by an LLM (involving fact checking, correction, restructure, etc) might be more complicated than just having someone write the original work.
When humans approach a piece of work like this, the production and verification processes pretty much happen at the same time, because the skills required to produce the work are the same ones that suggest the work is valid.
“Use a human” is often most clearly the sensible approach for projects that need a high level of accuracy and confidence in the material produced. For example, we talked to OpenFun about their LawTrace site, which brings together legislative information in Taiwan. They made a point of choosing not to use AI at all in this project. Having accurate information was far more important to users than any convenience AI could introduce.
4. Break down the verification problem
producing a solution is expensive / verifying the solution is expensiveSometimes solutions are expensive for a combination of reasons, and this can justify investment in trying to split the verification problem into smaller problems.
Through a sequence of different checks on LLM output, we can move problems towards being strong uses of AI, because it dramatically reduces the time needed to produce the solution, while the verification costs are manageable.
As an example, our APPG scraper sits in this category. We wanted to get accurate lists of parliamentary group memberships from dozens of different websites. Our original idea was that we would need to use a crowdsourcing approach, because we thought an LLM would be vulnerable to inventing lists of MPs.
But after some consideration, we invested time in a step where we could verify with code whether or the names extracted were actually listed on the relevant sites. We can see a similar example in the public consensus platform Pol.is – where category descriptions are linked back to concrete sources to facilitate easier double checking.
Similarly, you might find that aspects of your problem (if not the whole problem) are appropriate for mechanical checking. Could LLM code make a custom verification process easier? Can a series of automatic/human checks be made more efficient with a clear verification workflow? Each individual improvement moves your project closer to being a potentially strong use of AI.
Investment in the verification process might move the problem closer to having weak/strong AI benefits, where outputs can be derisked through cheap quality checks — but you’ll only know through systematically breaking it down in this way.
We hope that, by sharing this matrix, we will encourage more thoughtful deployments of AI technology in governments and beyond. Please feel free to share it with those who will find it useful.
—
This blog post has been adapted from our report Shifting Landscapes – A practical guide to pro-democratic tech.
-
In July 2024, we published our AI Framework, a first attempt at setting guidelines on how we utilise AI in our work at mySociety. The basic principles running through that framework can be summed up as “Use AI responsibly, and only when, all things considered, it is the best tool for the job”.
That position will serve us, and any organisation, for life — but, with such a fast-moving field and such speedy integration across so many of our areas of work, the consideration of ‘all things’ must happen at regular intervals. It’s important that we keep checking in, to ensure that the framework is still providing timely and relevant guidance that reflects the current circumstances.
With this in mind, we’ve implemented a three-pronged approach:
- The AI Framework itself: this basic set of principles and questions is designed to act as a resource that staff members can keep referring back to if in doubt — and is a living document which we regularly update to reflect these fast-moving times.
- Our AI register: an internal spreadsheet, where staff are required to note new uses of AI, so that we have an up to date picture of how it is being deployed in areas across the organisation. This has already collected diverse use cases, from generating the transcripts of our podcasts with third party AI-based tools, to our own experiments with LLMs, like detecting inappropriate use of WhatDoTheyKnow. Where we think there are learnings for others to benefit from, we’ve written up these use cases here on our blog.
- A regular ‘Generative AI and Machine Learning’ meeting: this gives us the opportunity to discuss entries on the register, ask questions and consider any issues that have arisen.
So, with all this analysis going on, what are the changes we’ve seen since the first iteration of the framework? As you might imagine, they’re far-reaching, both in our own practices and in the wider world.
In our development: There’s been increasing experimentation and use of coding agent tools from our developers, always with an eye to whether these are producing valid outputs that genuinely save time, or solve problems that other tools couldn’t.
Some of the datasets we create now rely on LLM processes, where flexibility in interpreting and transforming language helps us create and combine data in ways not possible through other methods. These include the collection of data on APPGs; and our WriteToThem Insights work.
We’ve experimented with some AI-based problem-solving on our websites and infrastructure: for example, screening for personal immigration requests that have mistakenly been submitted on WhatDoTheyKnow (a longstanding issue); and we are in the early stages of exploring machine learning approaches to help us understand potential ways of handling any abusive messages sent through WriteToThem.
As for the external landscape: increasingly, we’re seeing funding opportunities that centre around the responsible deployment of AI for the good of democracy or transparency; and in our international community, via our Communities of Practice and TICTeC, we’re hearing of both more and more innovative application of AI to civic tech tools; and ways of monitoring its use by the state.
—
Image: Steve Johnson
-
Most discussion and usage of LLMs is focused on high profile closed models such as OpenAI’s ChatGPT family, and Google’s Gemini – which are widely available and integrated into a range of existing products and services.
Because these are closed models, access and hosting of the models is controlled by the companies that create them. This presents a dilemma for civic tech organisations who believe in open source – where important parts of their processes can disappear into black boxes beyond your control. These may work well/be affordable today, but creates new risks. Specific models might become unavailable, there might be changes in pricing, and this represents lock-in to specific providers.
Open LLM models provide an alternative approach. In a familiar issue from open source licensing, there are different ways in which a model can be ‘open’. Open weights models have the final structure of the model released and can be run on your own hardware (Meta’s Llama model is an example of this). Fully open models have the underlying (open licenced) training data released, as well as the recipes and evaluation systems used in their training. AI2’s OLMo family of models and the recent Swiss AI institute’s Apertus model are examples of these. Somewhere in between these are approaches like IBM’s Granite models, where the model is released as open weights and the data was licensed to be able to train on (addressing copyright issues) but is not publicly accessible.
What are weights? Basically a model can be understood as a big network of connections – where the ‘weights’ are how strong (and influential) a connection is. What’s happening in the training process is a refinement of these weights as a result of being exposed to the training data. The weights at the end of the process are the trained model, and can be shared and used by others. But if you also have the training data and process, you can recreate the model step-by-step, with a clear audit trail of what’s in it.
Any kind of open weight model is practically appealing because they unlock new ways to work with private data without sharing with third parties, and create more flexibility around infrastructure. For instance, we currently use a fine-tuned version of Llama to help flag immigration correspondence in WhatDoTheyKnow.
Fully open models are ethically appealing because they avoid the issues of models that have been trained on copyrighted data. Their existence is a challenge to an AI policy debate where countries must trade-off the rights of creators against the benefits of AI as sold by a handful of companies. They fit well with our open source ethos – and understanding more about how to use them practically helps give us options to improve our own services, and contribute to wider arguments about responsible use of AI.
This blog post is a write-up of several practical experiments in using the 7b parameters variation of OLMo-2 both locally on a laptop GPU and remotely using HuggingFace’s inference endpoints.
Using OLMo-2 locally
Our purpose in running something locally is to be able to process sensitive information that should not leave our infrastructure. In this case, using OLMo-2 to create human-readable representations of clusters from WriteToThem survey responses. While users are asked not to include personal information in this survey, enough do that we need to treat the basic dataset as having personal information that should not be shared.
We used llama-cpp (and the associated python bindings) to run the local model. An alternative local approach is to use ollama to run a local server. The reason for using llama-cpp in this case is that ollama doesn’t always seem to pick up that less well known models can use ‘tools’ correctly (which is required for structured data output). Another benefit is having it run in process rather than as a separate server is the script can turn on and off the resource intensive bit (although there’s a corresponding start up time) rather than needing a separate server process to run.
Setting up the libraries
Installing llama-cpp in a way that can use the GPU is not straightforward. This set of instructions for Windows 11/Nvidia GPU mostly worked for me. I additionally needed to add an extra DLL directory before importing from llama_cpp because there’s a DLL folder that the library wasn’t yet referencing.
Big picture, WheelNext is a project to try and make installing correct versions of the library easier across different OS/GPU combinations. In the meantime, setting up a local machine is a bit fiddly.
Downloading model information
Llama-cpp uses GGFU files – which have all the weights in a single file. There are libraries to convert from the transformers format – but this is often made available by model publishers on HuggingFace.
Downloading the model can be done using the huggingface_hub command line too (here using uv).
uvx –from huggingface-hub hf download allenai/OLMo-2-1124-7B-Instruct-GGFU olmo-2-1124-7B-instruct-Q4_0.gguf –local-dir models
This is pulling down a quantised version – which has the same number of parameters – but the values of the weights have been significantly rounded down. This tends to have much less decrease in quality than the corresponding decrease in file/memory size (why? Broadly high fidelity here is useful for adjusting in training which will happen in small shifts, but when you have something working the general structure is good enough) – and this fits it just inside the ability of my laptop’s GPU.
This download can also just be done in code:
from llama_cpp import Llama
from functools import lru_cache
@lru_cache
def get_llm():
return Llama.from_pretrained(
repo_id=“allenai/OLMo-2-1124-7B-Instruct-GGUF”,
filename=“olmo-2-1124-7B-instruct-Q4_0.gguf”,
)
Structured data output
To get structured data out of the model, Pydantic AI can be used with Outline to query the llama cpp model.
This:
- makes it easier to define Pydantic data structures that should be returned.
- makes it easier to swap between local/remote models by swapping the model passed to the agent, but otherwise using a common API.
Hosted OLMo-2 model
An advantage of any open weights model is being able to run it on a range of infrastructure (and being able to change the infrastructure later).
In this case, I had a use case where we wanted to do transformations on already public data (the appropriateness of linking to a specific Wikipedia page from a specific sentence in a parliamentary debate) – and so there was no privacy/security issue for the purposes of the experiment. We are doing further exploration about how we can make this kind of use compliant with our wider legal and privacy commitments.
Because OLMo-2 is not a commonly used model, there isn’t an inference service that offers it directly as an option (which would be most efficient – as you’re being charged for tokens while the underlying infrastructure is shared between many users). Instead, you need to create a private server that can manage the model.
Creating an endpoint
Hugging Face Inference Endpoints is the approach I used here – that lets you provision an endpoint connected to a specific model. I’m using the same model as I used locally.
Depending on the properties of the model – the minimum GPU required will be suggested. This model was coming up about $0.8 an hour. Running the 13b parameter version of the model was about $2 an hour. There are options to run on AWS, Azure and Google Cloud in different regions (although processing data in the EU/UK is a requirement – this limits some of the GPU options).
The scale-to-zero time is adjustable down to about 15 minutes. It takes a few minutes to load up from this. In principle, if the access token is scoped correctly – the huggingface_hub library can handle pausing and unpausing the endpoint (or even programmatically creating one), if some more control here is wanted.
Structured data output
This endpoint works well using some of the example HuggingFace connections for PydanticAI. Something I had to adjust was adding an adapter to reduce complex json schemas (e.g. anything with multiple model types, enums, etc) from using ‘$defs’ to just being a normal structure because the Hugging Face text-generation-inference interface can’t handle them.
I have an example of creating a model that Pydantic AI will accept here – the missing config bits are a token associated with the account and the url of the endpoint created.
So in principle this means we can have an endpoint that gives us access to a GPU based model for an hour a day at a reasonable price – while we could at a later point swap out to use a local model without adjusting the general logic of the application. This is well suited to our current anticipated uses in batched backend processes, but would be less efficient if it needed to be responsive around the clock.
Reflecting on the results
Compared to previous projects using the OpenAI API, a key thing to note is it is slower and more fiddly on the infrastructure at hand. I was only using the 7b parameter model, while the 32b parameter model is the one that evaluates closer to GPT-4o mini. As such, prompts needed to be a bit more detailed on what was required. Similarly, a combination of the hardware and not being able to run queries in parallel over a wider infrastructure mean the process takes longer.
But this is also like comparing cake to a well balanced meal – the benefits of an open model are not just philosophical but practical. With a bit more work on the prompt you can get useful results on a laptop with no dependency on third-party services. That brings into scope a range of use cases that OpenAI is not suitable for.
Even where, such as in the Wikipedia example, there are no privacy issues in using OpenAI, making it easy to swap in an open model makes it much easier to evaluate the effect of using an open model. It will now be relatively straightforward to quickly substitute OLMo-2 into PydanticAI flows using other models and get a baseline feeling for effectiveness. Even where you might choose to use a closed model in a specific instance, it is very useful to work in such a way that you are not locked in to that model and could switch away in future.
Similarly, having a working process for a non-mainstream model like OLMo-2 makes it easier to explore other models like Apertus. As this has been trained on a wider range of non-English languages it could provide a more dependable component in LLM integration with the core Alavateli software – which powers Freedom of Information platforms across a range of languages.
Understanding open models as a practical approach helps contribute more widely to policy conversations around AI – and where trade-offs and impacts are inherent to the nature of the technology, or are a consequence of how they are currently controlled and produced.
Open models are always likely to lag slightly behind the frontier models, but they are already incredibly useful technologies compared to what was possible a few years ago. We want to understand more about how we can practically make use of these models – and help make sure the future of LLMs are shaped by ethical considerations about their training and use – rather than accepting them on the terms of the dominant tech giants.
Header image: Photo by Zhang Zi Han on Unsplash
-
Recently we wrote about why we’re now listing APPGs in TheyWorkForYou. This blog post goes into more detail about the technical process we use to gather who is a member of an APPG.
We have two methods of getting the memberships of APPGs. The first is finding if it’s already published on their website. The second is using Parliament’s rules to ask the APPG contact for the list. So we need to a) find all the APPG websites, and b) see if they publish members lists c) if not, ask for the list and d) get those lists into a consistent format.
Data that is fragmented and not in the format we want is a fairly common civic tech problem. The solution is to write a ‘scraper’ that reads the content of a website and has a process for converting it to a more structured format.
This works well when dealing with only a few sources (e.g. the memberships of the UK’s parliaments only needs a few different scrapers), or where a common format is being used (e.g. many local government websites use similar providers). In the case of APPGs, there is no common template being used. We just have a set of a few hundred websites that may (or may not) contain a list of names.
Rather than a traditional scraper, we have built an agentic AI/LLM approach that is more flexibly able to extract memberships from websites. The end result is a tool with a careful sequencing of manual and automated steps, injecting human review in structured ways. Rather than an “AI makes mistakes” disclaimer, we built a structured process to check elements efficiently one group at a time, that can lock off errors before proceeding to the next stage. This was also an experiment in using LLMs to write scraper tools, as well as some of the tools needed for the manual review steps.
Practically, this was an effective way of getting the information we needed that turned a very hard problem into one that we can dependably run regularly. It also suggests more generally useful ways of approaching fragmented data problems (more on this at the end of the post).
Building agentic approaches
An ‘agent’ is often poorly defined, but broadly it’s a language model interface is given tools (specific functions), a task, and an output data structure, and it loops between these until it gives a result.
To build agentic functions, we used the PydanticAI framework, which acts as a connector between the prompt, input data, the data structure of the output data, functions the agent has access to, and any bespoke validation of the results. The end result is a function that accepts structured input, and returns structured output, relatively painlessly.
Although this example is using OpenAI’s GPT models, in future experiments we use the PydanticAI approach to connect to open source models (the framework is designed to be model-agnostic). In principle this means that this project could in future switch the underlying provider used.
Process
Step 1: Writing a scraper
The first thing we needed to do was to get the official data from Parliament’s APPG register into a more structured form.
You can see an example of this page for the Africa APPG. This is a good task for a traditional scraper, but would also have been a fiddly problem. Using ChatGPT, we gave it an extract of the HTML, and asked for a Pydantic data structure and script to convert the data. This worked pretty well, with some tweaking to the format over time. When errors emerged in different APPGs – passing the error and an understanding of what should have happened back to the Copilot agent (using a Claude model) led to working fixes. In using the coding agent the key decision was deciding which bit of the project to be opinionated about – and this has mostly meant being very explicit about data structures (and validation to ensure they’re correct), and more relaxed about the pipes that connect things up.
Step 2: Adding categories to APPGs
From the official data, we only know if an APPG is a county or subject area group. We want to make it a bit more explorable by breaking this down into categories.
In the spirit of experimenting with LLMs, we copied all subject areas APPGs names and purpose statements into one of OpenAI’s reasoning models and asked for 10-20 sub-categories. It came back with 20 and they looked reasonable.
We then created a small functionless agent interface, giving it the title and purpose of a specific APPG, and returning a list of potential categories (preferring one, but allowing all that seem relevant).
Spot-checking these, they seem reasonable and for the purpose of breaking down the big list a bit – this is a good step up. This means, we can quickly see the APPGs that are likely to be relevant to environmental matters.
Step 3: Finding missing websites
Some APPGs list their external website – some do not. Here we use AI tools as part of the workflow, to find those missing sites (which may not exist).
We created an agent function with access to a web search tool (tavity), a function to check if the URL is valid, and a prompt to help identify the correct site. This creates a loop to search and identify a good candidate for the website.
At this point, there is a manual check that prompts the user to review each site one-by-one before confirming it as a valid site. 45/74 sites identified in the first wave were valid. Invalid websites were news articles, APPGs in other parliaments, or sites for previous iterations of that APPG.
This is not comprehensive and we and our volunteers found some more manually after the fact – but it is an interesting trial in finding data starting only with a search engine.
Step 4: Find published members
The final step is to get a list of members (if published) off these websites. We need a really flexible approach for this. Names might be in a structured list, but they can also be in one paragraph. They might be on a members page, the home page, or spread over three pages. There is no consistency to fall back on.
Here, we created an agent with a function that can fetch a web page and convert it to markdown. Using this recursively, the prompt instructs the agent to find the most relevant page (in some cases pages) that could contain membership information, and return a data structure of the members (MPs, Lords, Other). This returned over 5,000 names in the data format provided.
The big risk at this point is that having been asked for a list of MPs, it makes some up. The validation we use for this is to check if each name in the list is present within the HTML content of the page it was extracted from. If there’s an error, it runs again and will give up rather than use an incorrect list. There is some possibility for misinterpretation – but this prevents outright fabrication. Errors flagged here tended to be when the LLM has fixed formatting meaning the text no longer matches exactly against the page.
The key problem here is one that a human would have too – some APPG lists are out of date. Here I added an extra flag detecting a list containing people who had left Parliament that then needed a manual review. In other cases, this was sometimes picking up lists that were not membership lists. We made some adjustments to the prompt after picking up attendees at the AGM – which is not wrong, but incomplete.
Step 5. Manual data
As our main blog post talks about, we then needed to contact APPGs directly for lists that were not published. This presented a new problem: what we got back was a combination of spreadsheets and emails with different levels of detail – some including party details in other columns, some not.
Our solution was to have a Google Doc that just has each list formatted under a heading with the APPG title – we could just copy and paste information into this.
This file is then downloaded as markdown and converted into a list of names. There are a few tweaks to clean up leading numbers, and identify the name component of the line. Again, this step was substantially written via prompt – giving the LLM examples of the problem data, and that would create regular expressions to clean the data into the basic list of names we needed.
Step 6: Tidy members information
What we want to do next is get from a list of names to a list of TheyWorkForYou unique IDs.
We have a library that helps reconcile names to IDs, but a challenge here is that there are a huge range of spelling mistakes (sometimes to an extent where you could not actually work out the correct MP).
What we needed was a quick tool to compare the input name against our list of known names and suggest near matches. Here we again turned to the coding agent, posing the problem, providing some snippets to interact with our existing library, and letting it craft a command line interface.
This fairly quickly gave a good interface for reviewing spelling problems (which was later refined to auto-match below a certain threshold). This helper tool is not especially complicated, but as something with a clear input and output, isolated from the rest of the flow, was a good candidate for testing using Copilot to create the function. In choosing what to spend time on, this would not otherwise have been a priority – but brought a useful feature into scope.
Result
The end result of this process is fairly effective – with a series of steps we can repeat every six weeks when a new APPG register is released to check for new webpages for new APPGs, or to recheck previously scanned pages.
The efficient sequencing of steps means that manual review happens on similar tasks in sequence, rather than checking each APPG through all steps.
In general, I’m pretty happy with the results of this, it made a project that would otherwise only have been possible with a big (and fairly boring for participants) crowdsourcing effort possible.
One of the problems we have to deal with a lot is fragmented public data, when relevant data is scattered all over the place and is a lot of work to bring back together. Here we found AI tools that were both useful in discovery of a component of the data, and in reconciling to a common standard.
The “AI scrapes then verifies content is present” approach worked well here but would struggle with more complex problems. For instance, if we really needed to be sure we were extracting a correct party label alongside a name, knowing that ‘Labour’ was present on the page wouldn’t be as helpful.
Building on this, the AI-written scraper code worked pretty well. If properly sandboxed (pydantic-ai has support for running python in a sandbox using pyodide), transformation code could be written to convert data between different sets of headers without running the data itself through an LLM to convert it. This potentially helps with some of the fragmented data problems of reconciling compatible but different schemas. LLM-involved approaches have a real potential to create new datasets through easier discovery and joining of data.
This is a way we can use new technology to make a dataset possible, but also it would be much easier if Parliament gathered and published this in the first place. The equivalent Cross Party Groups in the Scottish Parliament just make a downloadable file of all memberships in their open data portal. We need to think about how new technological approaches are not just propping up bad transparency – but part of encouraging better transparency all the way upstream.
Header image: Photo by Susan Holt Simpson on Unsplash
-
The government is making a significant investment into AI in public services, and systems are changing apace.
AI is increasingly being deployed in every department of government, both national and local, and often through systems procured from external contractors.
In a recent article for Public Technology, mySociety’s Chief Executive Louise Crow flags that we urgently need to update our transparency and accountability mechanisms to keep pace with the automation of state decision-making.
This rapid adoption needs scrutiny: not only because significant amounts of money are being spent; but also because we’re looking at a new generation of digital systems in which the rules of operation are, by their very nature, opaque.
To see Louise’s thoughts on what needs to change, and why, as this new technological era unfolds, read the full piece here.
If you find it of interest, you may also wish to watch this recent event at the Institute for Government, The Freedom of Information Act at 25, where Louise was one of six speakers reflecting on the future of transparency in the UK.
—
Image: Alex Socra
-
mySociety was founded on one seismic technological change: the arrival of the internet, bringing radical new possibilities to the ways in which we engage with democracy.
Now we’re seeing a second upheaval, just as potentially explosive: the wide adoption of generative AI and machine learning tools — particular kinds of artificial intelligence — not least by the UK government, who have made a commitment to see AI “mainlined into the veins of the nation”.
From the visible and novel, like ‘AI bot’ MPs; to the hidden and less-interrogated, like the algorithms that drive decision-making around benefits; to the new capabilities around working with large text datasets that we ourselves are experimenting with at mySociety: artificial intelligence is changing the way democracy works.
We’ve been thinking about AI for some time, as have our colleagues around the world — TICTeC 2025 had a strong strand of pro-democracy organisations showcasing how they are using new technologies to hold authorities to account and support public engagement; alongside developers showing the tools that aim to make the government more responsive.
AI is coming to democracy, whether we like it or not. In many places, it’s already here.
But there are implementations in which it can be highly beneficial to us all; and ways in which it can present a clear and present danger to democracy.
It benefits everyone if there is a high level of understanding of both the challenges and the opportunities of AI in government. Democratic decision makers need to understand digital tech in order to legislate effectively around it, to develop and procure it effectively.
This is not just so that they can deliver services more efficiently, but also to ensure that they retain the legitimacy of democratic government by using tech and AI in a way that ensures transparency and accountability, preserves public trust and allows the public to understand and participate in the decisions that affect their lives.
Reflections for our time
Over the next few months, we’ll be sharing our own thoughts and experience — alongside invited guest writers who are thinking about how AI interacts with democratic processes and institutions, and how to make that better — in a series of short pieces.
These will examine the different ways that AI is affecting the things we care about here at mySociety:
- Transparent, informed, responsive democratic institutions
- Politicians and public servants who work for the public interest
- Democratic equality for citizens: equal access to information, representation and voice
- A flourishing civil society ecosystem
- The effective and principled use of digital technologies
- Action from politicians to match the evidence of the climate crisis and the level of public concern
- Better communication between politicians and the public, creating space for climate action.
Stay informed
If you’d like to get updates in your inbox, make sure you’ve checked ‘artificial intelligence’ as an interest on our newsletter sign-up form (if you already receive our newsletters, don’t worry – so long as you use the same email address, this will just update your preferences. Just make sure you’ve ticked everything you’re interested in).
By also completing the ‘how do you identify yourself’ section, you’ll help us send you the most relevant material: that means guidance if you work in government or build tech; data and our analysis if you’re a researcher; tools for holding authorities to account if you are an individual or work in civil society, and so on.
—
Image: Adi Goldstein