Singularity AI

Research

The four threads

These are the areas we read most closely. They are not separate fields so much as four views of the same difficulty: we can build systems whose behaviour we cannot fully specify, predict or explain.

01

Alignment

The gap between the objective a system is trained on and the intention behind it. Reward modelling, preference learning, specification gaming, and the awkward fact that a proxy optimised hard enough stops behaving like the thing it proxied.

02

Interpretability

Mechanistic work on features and circuits, sparse autoencoders, attribution methods, and the standing question of faithfulness — whether an explanation describes the computation that actually occurred or merely one consistent with the output.

03

Safety evaluation

How capability and risk are measured: benchmark construction, contamination, red-teaming methodology, jailbreak taxonomies, and the reliability gap between an evaluation score and deployed behaviour under adversarial pressure.

04

Governance

The EU AI Act's tiering, the UK's sector-led posture, model release norms, compute thresholds as a regulatory handle, and what voluntary commitments have and have not achieved in practice.

How we read a paper

A rough working method, offered because too much AI commentary skips it.

  • Claim before result. What exactly is being asserted, and over what population?
  • Baseline discipline. Compared against what, tuned how hard, and by whom?
  • Contamination check. Could the evaluation set plausibly appear in training data?
  • Effect size, not significance. Does the improvement matter at the scale it claims to?
  • Negative space. Which obvious experiment is missing, and would it have been cheap to run?

None of this is exotic. It is ordinary scientific reading, applied to a field where the publication cycle is fast enough that it often gets skipped.