01
Alignment
The gap between the objective a system is trained on and the intention behind it. Reward
modelling, preference learning, specification gaming, and the awkward fact that a proxy
optimised hard enough stops behaving like the thing it proxied.
02
Interpretability
Mechanistic work on features and circuits, sparse autoencoders, attribution methods,
and the standing question of faithfulness — whether an explanation describes the computation
that actually occurred or merely one consistent with the output.
03
Safety evaluation
How capability and risk are measured: benchmark construction, contamination, red-teaming
methodology, jailbreak taxonomies, and the reliability gap between an evaluation score and
deployed behaviour under adversarial pressure.
04
Governance
The EU AI Act's tiering, the UK's sector-led posture, model release norms, compute
thresholds as a regulatory handle, and what voluntary commitments have and have not
achieved in practice.