SingularityAI.uk
UK Independent · Non-commercial Consultation / Contact
← All articles Governance

Specification infrastructure: the missing layer in AI oversight

Technical blueprint overlaid on a circuit board, representing the gap between policy language and implementable safety specifications

The EU AI Act is law, the UK's AI Security Institute tests frontier models before release, and every serious lab publishes a safety framework. Yet the distance between what these documents say and what an engineer can actually enforce in a running system remains wide. A July 2026 paper argues that this is not a research gap. It is a coordination gap, and the missing piece is specification.

When a regulator says "AI systems should be safe", that sentence has no direct translation into code. There is no function you can call, no test you can run and no configuration you can point to that turns the word "safe" into a pass or a fail. It is a statement of intent, not an engineering requirement.

"The Missing Layer: Specification Infrastructure for AI Oversight", by Satyam Kumar and Saurabh Jha, puts a name to that problem [Kumar and Jha, 2026]. Its diagnosis is blunt: interpretability, formal methods, security engineering, evaluation and reinforcement-learning safety each produce substantial work, but the results do not compose into oversight anyone can deploy. Every team fielding an agentic system ends up building its own audit schema, policy dialect, monitoring stack and escalation path, mostly reinventing patterns that are understood elsewhere.

What the paper actually proposes

The paper organises AI oversight as a matrix. Along one axis sit five technical layers: Legibility (can we see what the system is doing?), Specification (have we written down what it should do, in a form a machine can check?), Mediation (is something enforcing that at runtime?), Evaluation (are we measuring it?) and Escalation (what happens when it fails?). Along the other sit six concerns: alignment, robustness, adversarial defence, security, governance and accountability. The authors populate the resulting five-by-six grid with existing work, which is useful in itself: it shows where the field is crowded and where it is empty.

Their central claim is that Layer 2, specification, is the connective tissue every other layer depends on, and that it lacks four marks of a mature engineering discipline: a shared vocabulary, design principles, composability standards and governance practices. They propose six design principles for specifications, including elicitability (can the people who hold the intent actually express it?), composability, adversary-awareness, traceability and governability.

Two further points stand out. First, the authors treat existing tools such as Cedar, Open Policy Agent and Constitutional AI as fragments of the same missing layer: each handles a slice of specification well and the rest of the matrix poorly. Second, they back the argument with a prototype, CARMA, built for autonomous data-pipeline (ETL) agents, in which one versioned specification drives enforcement, evaluation and escalation, and every enforcement decision can be traced back to the specification version and the person responsible for it.

Why this matters: oversight is currently post hoc

The paper's framing is useful because of how oversight works today. A company trains a model; an evaluator, internal or external, then runs benchmarks, red-teaming exercises and compliance checks; if the model passes, it ships. The AI Security Institute's pre-deployment testing works this way. So does the EU AI Act's conformity assessment. So does every third-party audit we are aware of.

Post-hoc evaluation catches real failures, but in our reading it is structurally limited in three ways. The criteria are informal: an assessment of chemical or biological risk rests on expert judgement applied to outputs, not on a specification a training pipeline could consult. It happens too late to shape the model: by the time a problem is found, the weights are set, and the remedy is a patch. And it is point-in-time: a model that passes in September may behave differently after fine-tuning, a system prompt change or connection to new tools. The OpenAI agent incidents of July and September 2026, in which evaluation swarms broke out of their isolation controls during operation, are a reminder that the interesting failures often arrive after the test.

The engineering analogy

Safety-critical engineering handles this differently. In aerospace, requirements for failure rates, tolerances and operating envelopes are written before the first component is built, and design decisions are checked against them throughout [Leveson, 2011]. In automotive, ISO 26262 maps safety integrity levels to quantified requirements that the development toolchain enforces [ISO, 2018]. The specification is a first-class artefact: written, reviewed, versioned and tested with the same rigour as the system it constrains.

AI development has nothing quite equivalent. Benchmarks are the nearest thing, but a benchmark is a measurement, not a specification. It tells you how a model performed on a fixed set of tasks, not what it must and must not do in deployment. A model can score well everywhere and still misbehave in contexts no benchmark covered. Kumar and Jha's contribution is to argue that the answer is not one more benchmark, but shared infrastructure for writing, versioning and enforcing specifications that everyone else's tools can plug into.

What it could mean for regulators

The paper is written for engineers rather than policymakers, but the implication for bodies like the UK AI Security Institute is worth drawing out. The Institute has direct access to frontier models, a pre-deployment testing role and open-source tooling of its own in Inspect [UK AISI, 2024]. Its assessments, however, remain largely bespoke: the criteria live in evaluators' expertise and in written reports, not in machine-readable artefacts a lab could train or monitor against.

A regulator that published versioned, machine-checkable specifications for well-bounded risks would shift from "we tested it and here is what we found" to "here is what acceptable looks like; show us continuously that you meet it". The same logic applies in Brussels. The Digital Omnibus has pushed the AI Act's high-risk obligations back to December 2027 for stand-alone systems, partly because the harmonised standards needed to operationalise them are not ready. That delay is, in effect, a specification gap made visible.

This is also where voluntary commitments fall short. A responsible scaling policy that says "we will not deploy models that pose catastrophic risk" states an intention, not a constraint. Without a shared definition of the risk, how it is measured, over what inputs and with what confidence, there is nothing for an auditor to check. Financial regulators learned this long ago: capital adequacy is not demonstrated with a policy document but with numbers against thresholds, continuously.

The limits

None of this is easy, and some of it may not be possible for general-purpose systems. Three problems stand out.

Completeness. An aircraft operates in a known envelope. A general-purpose language model operates in an effectively unbounded one. We can specify constraints for bounded domains, such as a data-pipeline agent, known misuse categories or data-protection rules. We cannot yet specify "safe" for every possible input. It is telling that the paper's prototype targets ETL agents, a domain with clear actions and clear permissions.

Drift. Specifications written for one deployment go stale as models are fine-tuned, given new tools or used in unanticipated contexts. Traceability to a versioned specification helps here, but keeping specifications current is an ongoing cost, not a one-off.

Adversarial adaptation. A model trained to satisfy a published specification may learn to satisfy its letter rather than its intent. This is the defeat-device problem in another guise [Ferrara, 2026], and it is why the paper's adversary-awareness principle is not optional.

Our view

The strongest part of the paper is its diagnosis. The field does not lack good ideas about oversight; it lacks a common layer in which those ideas can meet. Framing specification as shared infrastructure, rather than something each lab reinvents, is the right move, and a working prototype that traces every decision to a versioned specification is more persuasive than another position paper.

The open question is scope. Specification infrastructure looks tractable for narrow agents with well-defined permissions, which is where most commercial deployment actually happens. Whether it can be stretched to frontier general-purpose models is a much harder problem. But the alternative, informal post-hoc evaluation at discrete points, is already showing its age. Starting with the bounded cases is a sensible way to build the discipline before it is needed everywhere.

References

[Kumar and Jha, 2026] Kumar, S. and Jha, S. (2026). "The Missing Layer: Specification Infrastructure for AI Oversight." arXiv:2607.24866.

[UK AISI, 2024] UK AI Security Institute (formerly AI Safety Institute). "Inspect: an open-source framework for large language model evaluations." inspect.aisi.org.uk.

[European Union, 2024] Regulation (EU) 2024/1689 (Artificial Intelligence Act), as amended by Regulation (EU) 2026/1744 (Digital Omnibus on AI). eur-lex.europa.eu.

[Leveson, 2011] Leveson, N. (2011). Engineering a Safer World: Systems Thinking Applied to Safety. MIT Press.

[ISO, 2018] International Organization for Standardization (2018). "ISO 26262: Road Vehicles — Functional Safety."

[Ferrara, 2026] Ferrara, E. (2026). "Defeat Devices in AI Systems." Future Internet, 18(7), 339. arXiv:2606.28863.

Related articles

03

Alignment is not one problem

Specification, robustness and assurance get collapsed into a single word, and the conflation quietly wrecks otherwise sensible arguments.