Making systems want what we want
Reward modelling, RLHF and its successors, specification gaming, and the persistent gap between a training objective and an intention.
Independent · Non-commercial · UK
Singularity AI is an independent commentary publication covering alignment, interpretability, safety and governance. We read the papers, check the claims against the evidence, and write up what actually holds — including when the honest answer is that nobody knows yet.
Four threads run through everything here. They overlap constantly, which is rather the point.
Reward modelling, RLHF and its successors, specification gaming, and the persistent gap between a training objective and an intention.
Mechanistic interpretability, feature attribution, sparse autoencoders, and the question of whether an explanation is faithful or merely plausible.
Evaluation design, red-teaming, jailbreak taxonomies, and why benchmark scores routinely overstate real-world reliability.
The EU AI Act, UK regulatory posture, model release norms, compute governance, and the practical limits of voluntary commitments.
Longer-form pieces, published when there is something worth saying.
Saturation, contamination and construct validity — why a leaderboard number tells you less about capability than it appears to.
Sparse autoencoders gave the field real traction. They did not give it the ability to certify a model as safe, and the distinction matters.
Specification, robustness and assurance get collapsed into a single word, and the conflation quietly wrecks otherwise sensible arguments.
A note on independence. Singularity AI takes no funding from AI laboratories, vendors or advocacy organisations. Where a piece discusses a system we have commercial exposure to, that is disclosed in the article itself.