Architectures of Inference

Institutions, Assurance, and What Happens When Models Become Assessors

AI
Governance
Psychometric validation
From a lecture on Britain’s AI future to the psychometrics of LLM raters — on why the harder problem is not building capable models but building the institutions that can judge what their outputs actually mean
Published

July 29, 2026

Artificial intelligence is forcing us to ask not only what machines will be capable of, but what kinds of institutions we will need once those capabilities become ordinary. For years now, public discussion has centred on model performance. Can an AI system write, reason, code, diagnose disease, discover drugs or pass professional examinations? But as these capabilities move from demonstrations into workplaces, schools and public services, the more consequential questions concern what happens around the model.

How should people be educated in a world where intelligent assistance is ubiquitous? How should organisations decide when an AI system is trustworthy? What evidence should support decisions made with, or delegated to, these systems?

A recent discussion between Sir Demis Hassabis and Dame Wendy Hall provides one route into these questions. The UK Government’s work on AI assurance provides another. And a growing body of research into AI-based scoring shows why the scientific problem may be harder than simply producing another benchmark. Taken together, they point towards the emergence of something larger than any individual model: an architecture through which machine-generated inferences can be produced, evaluated and justified.

Act I: The horse has bolted

In the Worshipful Company of Information Technologists’ annual lecture, Demis Hassabis and Wendy Hall discussed the future of artificial intelligence, Britain’s position within it and the changes required across science, education and government. The conversation ranged widely, but its most important subject was not the technical architecture of AI systems. It was the institutional architecture that will surround them.1

On education, Hassabis’ position was clear: the horse has bolted. Students are already using generative AI, and attempting to preserve the previous educational settlement by simply prohibiting it is unlikely to work.

Instead, he proposed something closer to an inverted classroom or “Montessori Plus” model. AI tutors could support personalised instruction, factual learning and practice, while time spent with teachers and other students could concentrate on the things that are harder to automate: project work, collaboration, creativity, entrepreneurship, judgement and interpersonal capability.

The argument is not that knowledge no longer matters. It is that the relationship between knowledge and capability is changing.

Where information is abundant, locating an answer becomes less important than interpreting it. Where fluent content can be produced instantly, greater weight falls on deciding what should be produced, whether it is correct and what ought to follow from it. The educational task shifts from reproducing knowledge towards navigating, testing and applying it.

Wendy Hall’s contribution placed this challenge within a longer institutional history. She has been involved in the development of the UK’s national AI capability across several stages, including the 2017 independent review of the UK AI industry and the institutions that followed. Her emphasis was not simply that individuals need to learn new skills, but that universities, governments and professional bodies must develop the capacity to respond coherently.

That distinction matters. “AI literacy” can easily be reduced to teaching people how to prompt a chatbot. Institutional capability requires something more demanding: expertise, standards, curricula, professional roles, public infrastructure and a shared language for understanding what these systems do.

Hassabis offered what amounted to a hierarchy of challenges.

At the first level sit the technical questions. Can we understand, test and control increasingly capable systems? Can we establish meaningful evaluations and guardrails?

Beyond those lie questions of political economy. Who owns the systems? Who benefits from the productivity they generate? What happens to labour markets, wages and the distribution of wealth?

Finally come questions of meaning. If intelligent machines alter the relationship between labour, scarcity and human achievement, what becomes the basis of purpose and status?

Those higher-order questions are difficult to answer while the underlying institutions remain underdeveloped. Before a society can decide how it wishes to live with advanced AI, it needs the capacity to understand what the systems are doing and to intervene when necessary.

The future of AI is therefore not only a problem of technological development.

It is a problem of institutional development.

Act II: From governance to assurance

One attempt to build that institutional capacity can be found in the UK Government’s Introduction to AI assurance, published by the Department for Science, Innovation and Technology in February 2024.2 It is deliberately introductory — the first in a promised series — and it inherits its framing from the 2023 white paper A pro-innovation approach to AI regulation, which set out five cross-sectoral principles for the responsible use of AI: safety, security and robustness; appropriate transparency and explainability; fairness; accountability and governance; and contestability and redress.

Those principles describe outcomes. Assurance is what connects them to evidence. The distinction can be expressed simply:

Governance describes what we expect. Assurance provides evidence about whether it is happening.

The guidance defines assurance as the process of measuring, evaluating and communicating something about a system, process, product or organisation — in the case of AI, its trustworthiness. The term is not new. It is borrowed from accountancy and has been adapted before, for cyber security and quality management, fields the document returns to as its model for what a mature assurance ecosystem looks like. (The UK’s cyber security industry, it notes, is already worth close to £4 billion.)

Each part of that definition carries weight.

Measurement involves gathering qualitative and quantitative information about how a system behaves: its performance, functionality, limitations and potential effects in different contexts.

Evaluation involves interpreting that evidence against standards, benchmarks, requirements and intended uses.

Communication makes the resulting evidence available to the people who must act on it, whether they are developers, senior leaders, regulators, purchasers, frontline workers or members of the public.

The most useful idea in the document is a distinction it draws between three things that are easily conflated. Trust is whether a person or group is willing to rely on an AI system. Trustworthiness is whether the system actually deserves that reliance, on the basis of reliable evidence. Justified trust is the overlap: trust that is warranted because the evidence supports it.

The gap between the first two is where the problems live. Trust without trustworthiness is misplaced confidence — a system relied upon that has not earned it. Trustworthiness without trust is wasted potential — a sound system no one is willing to use. Assurance is the machinery for closing that gap in both directions: producing the evidence that lets a trustworthy system be trusted, and withholding it from one that should not be.

This reframes trust as something other than confidence in a brand, a model provider or the general promise of technological progress. It becomes a conclusion drawn from reliable, standardised and accessible evidence that a system works as intended, that its limitations are understood and that its risks are being managed.

The guidance does not present a single test that can establish trustworthiness. It describes a toolbox.

Depending on the system and its context, assurance might involve performance testing, conformity assessment, impact assessment, risk assessment, bias audit, formal verification, documentation, monitoring, certification or independent inspection. These mechanisms may be applied at different stages of the AI lifecycle and combined according to the nature of the claim being made.

This is one of the document’s strengths. There is no universal assurance score that can tell us whether an AI system is “safe” or “fair.” Evidence must be connected to the system, its users, its environment and the consequences of error.

A model used to recommend films does not require the same evidence as one used to allocate medical treatment. A system that drafts internal notes is not equivalent to one that scores candidates for employment. The technical component may be similar, but the inferential and social contexts are not.

The programme also treats the surrounding ecosystem as an economic and professional domain in its own right: standards bodies, auditors, evaluators, technical specialists and regulators, alongside the organisations that commission and interpret their work. This is the institutional demand identified by Hassabis and Hall, now seen from the inside — and, as they argued, one that model developers alone cannot satisfy.

Yet, being an introduction, the guidance stops short of the hardest part. It explains why organisations require evidence and outlines the mechanisms through which that evidence might be generated. It does not fully resolve the more difficult scientific question underneath them:

What makes the evidence good enough to justify a particular conclusion?

That problem becomes especially visible when AI systems stop merely producing content and begin producing measurements.

Act III: When models become assessors

Consider an employment interview.

A candidate provides a series of answers. An assessor observes the candidate’s behaviour, identifies evidence relevant to particular competencies and converts those observations into scores. Those scores may then contribute to a decision about whether the candidate progresses or receives an offer.

Even in a conventional interview, the decision depends on a chain of inference.

Does the response demonstrate the intended competency? Did the assessor interpret it consistently? Does the scoring system distinguish relevant differences between candidates? Does performance in the interview predict anything important outside it? Are the conclusions equally defensible across demographic groups?

Psychometric practice developed in large part to evaluate chains of inference such as these. Tests and assessments are not considered trustworthy merely because they produce numbers. Their use must be supported by evidence concerning reliability, construct representation, relationships with other variables, fairness and the consequences of interpretation.

These questions are now being transferred to AI systems.

A recent study by Kayden Stockdale, Louis Hickman and Siyi Liu put this transfer to the test.3 Asking whether large language models can serve as alternative raters of employment interviews — systems that evaluate open-ended responses against behavioural criteria, personality constructs or competency frameworks without a separately trained model for each scoring task — they scored two datasets and then subjected the resulting scores to the full apparatus of psychometric validation: intrarater reliability, test–retest stability, convergent, discriminant and criterion evidence, group differences and measurement bias, benchmarked wherever possible against human raters and against earlier supervised machine-learning models.

The attraction is obvious. Human scoring is expensive and time-consuming. Assessors disagree, become fatigued and require training. A language model can apply the same instructions across thousands of responses and generate scores almost instantly.

But consistency of application is not the same as validity of inference.

A model may produce highly repeatable scores while systematically attending to the wrong features. It may agree with human raters because it reproduces their shared biases. Its apparent accuracy may change with the wording of the prompt, the model version, the temperature setting, the number of examples provided or the order in which criteria are presented.

Stockdale and colleagues found exactly this pattern. Choices that look like mere configuration — model size and version, temperature, how much construct detail the prompt carries, whether candidates are scored singly or in batches, whether several runs are pooled into an ensemble — proved psychometrically consequential, and their benefits could plateau or even reverse, echoing a long-known result in human interviewing where added structure raises validity only to a ceiling. Their verdict was carefully hedged: ensembles of larger, newer models given detailed construct information could match or exceed supervised machine-learning models and single human raters — yet they still urged caution before such systems are trusted with high-stakes decisions, precisely because the evidence needed to justify that trust is only beginning to accumulate.

The introduction of an AI rater does not remove the measurement problem. It adds new components to it.

If the interview study shows what validating an AI rater demands, automated essay scoring shows how quickly the underlying capability is moving.

Earlier automated scoring systems often predicted a single overall score. Multi-trait assessment — separately evaluating content, organisation, fluency, word choice, conventions and other dimensions — typically required multiple output layers or separate trait-specific models.

The autoregressive multi-trait scoring approach proposed by Do, Kim and Lee4 treats the task differently. Rather than framing essay assessment as ordinary regression or classification, it reframes scoring as sequence generation. A single language model generates several trait scores as an ordered sequence, with each later prediction able to condition on the scores already produced.

This allows relationships among traits to become part of the inferential process. A judgement about overall quality need not be generated independently of judgements about content or organisation. The scores become an interconnected profile rather than a collection of isolated predictions.

In experiments using the ASAP and ASAP++ datasets, the authors reported average improvements of more than five per cent across prompts and traits relative to their baseline. The approach also allowed one model to generate multiple scores across several prompts, avoiding the need to duplicate large scoring architectures for each criterion.

The paper is careful about its limitations. Performance deteriorated for traits with very small amounts of training data. The ordering of score generation remained an open question. The authors also noted that further work was needed to explore alternative pretrained models and the conditions under which autoregressive scoring works best.

Those limitations are not incidental. They show why an assurance framework must extend beyond a headline performance result.

Changing the order in which traits are scored may change the resulting profile. Data scarcity may affect some dimensions more than others. A model may show strong average agreement while remaining unreliable for particular prompts, groups or score ranges. Efficiency gains may make a system easier to deploy while also concentrating several consequential judgements inside one opaque inferential process.

The engineering question is whether a model can generate the scores.

The validation question is what those scores allow us to conclude.

This distinction is essential as AI moves into recruitment, education, healthcare and other settings where outputs are interpreted as evidence about people. A fluent explanation, a candidate rating and a competency profile may all be generated as text, but they occupy different epistemic roles.

One is content.

The other is measurement.

Once a model’s output becomes a measurement, its quality cannot be established by surface plausibility alone. Nor is agreement with a human rater sufficient. We need to ask what construct the score represents, how stable it is, how it relates to external criteria, where it fails and whether its intended use is justified.

This is where psychometric validation and AI assurance begin to converge.

AI assurance supplies the organisational requirement: systems should be measured, evaluated and their trustworthiness communicated.

Validation supplies the evidential discipline: the claims made from those measurements must be supported by an explicit argument and an appropriate body of evidence.

The resulting architecture includes the model, but it does not end there. It also includes the scoring framework, data, prompts, benchmarks, monitoring processes, standards, documentation, human oversight and the institutions responsible for deciding whether the resulting evidence is adequate.

These are architectures of inference: structures through which observations are transformed into conclusions, and through which those conclusions acquire sufficient legitimacy to be acted upon.

Towards a Culture of our own

Near the end of the discussion, Hassabis pointed towards Iain M. Banks’ Culture novels as a useful vision of a world after the development of advanced artificial intelligence.

It is a reference he has made elsewhere. Hassabis has described the series as an optimistic depiction of a post-AGI future: a civilisation in which humans coexist with enormously capable machine intelligences and are able to flourish alongside them.5

The contrast with another familiar science-fiction response is instructive.

In Dune, the Butlerian Jihad leads to a prohibition against thinking machines. The lesson often drawn is that human freedom requires their rejection.

Banks imagines something different. The Culture has not protected humanity by preventing machine intelligence from emerging. It has developed a civilisation capable of living with it.

That does not make the Culture uncomplicated. Its machine Minds possess extraordinary power, and the novels repeatedly examine intervention, paternalism, freedom and the moral compromises of apparently benevolent systems. But that ambiguity is part of what makes the reference useful. The question is not simply whether advanced AI is good or bad. It is what political, cultural and institutional conditions would allow coexistence without surrendering human agency.

The Culture is optimistic not only because its machines are intelligent, but because its civilisation has become capable of accommodating that intelligence.

The three acts of the present moment may therefore be inseparable. We need educational and social institutions suited to a world in which intelligent assistance is widespread; an assurance ecosystem capable of producing credible evidence about the systems we deploy; and scientific practices of validation that can tell us when the outputs of those systems actually support the conclusions we draw from them.

Building more capable models is only one part of that task. The larger part is building institutions capable of understanding, evaluating and governing what the models do — which may be the real work required before we can begin moving towards a Culture of our own.

Footnotes

  1. The Future of AI with Sir Demis Hassabis and Dame Wendy Hall, the Worshipful Company of Information Technologists’ annual lecture. https://www.youtube.com/watch?v=HlLa5iA8lOs↩︎

  2. Department for Science, Innovation and Technology. Introduction to AI assurance (February 2024). https://www.gov.uk/government/publications/introduction-to-ai-assurance/introduction-to-ai-assurance↩︎

  3. Stockdale, K., Hickman, L., & Liu, S. (2026). Scoring employment interviews with large language models: Evaluation design components, validity investigations, and best practice recommendations. Journal of Applied Psychology. https://doi.org/10.1037/apl0001396↩︎

  4. Do, H., Kim, Y., & Lee, G. G. (2024). Autoregressive Score Generation for Multi-trait Essay Scoring. https://arxiv.org/abs/2403.08332↩︎

  5. Demis Hassabis on Culture by Iain M. Banks, Recommentions. https://recommentions.com/demis-hassabis/books/culture-by-iain-banks/↩︎