Making systems want what we want
Reward modelling, RLHF and its successors, specification gaming, and the persistent gap between a training objective and an intention.
Singularity AI is an independent commentary publication covering alignment, interpretability, safety and governance. We read the papers, check the claims against the evidence, and write up what actually holds — including when the honest answer is that nobody knows yet.
Four threads run through everything here. They overlap constantly, which is rather the point.
Reward modelling, RLHF and its successors, specification gaming, and the persistent gap between a training objective and an intention.
Mechanistic interpretability, feature attribution, sparse autoencoders, and the question of whether an explanation is faithful or merely plausible.
Evaluation design, red-teaming, jailbreak taxonomies, and why benchmark scores routinely overstate real-world reliability.
The EU AI Act, UK regulatory posture, model release norms, compute governance, and the practical limits of voluntary commitments.
Method
We work from the technical publications and system cards themselves — not press releases, launch posts or briefings.
Uncertainty is recorded, not smoothed over. When the evidence is inconclusive, we say that nobody knows yet.
No lab funding, vendor advertising or advocacy ties. Independence is the whole value of the exercise.
Singularity AI takes no funding from AI laboratories, vendors or advocacy organisations. Where a piece discusses a system we have commercial exposure to, that is disclosed in the article itself.
Dispatch archive
Longer-form pieces, published when there is something worth saying.
A Lean-checked proof of finite-time blow-up, a Millennium Prize problem apparently settled, and a credit dispute that matters as much as the maths.
Stuart Russell's case that safety requirements must come first, and why "pacing the frontier" gets the order backwards.
Sam Altman says so. We check the claim against the evidence, including where the much-quoted 52x figure really comes from.
Get in touch about a paper, a model card, a policy question or a correction to something we have written.