Articles
Everything we have published
Long-form commentary on alignment, interpretability, evaluation and governance. Published when there is something worth saying, which is less often than a content calendar would like.
-
Evaluation
Benchmarks mostly measure benchmarks
Saturation, contamination and construct validity — why a leaderboard number tells you less about capability than it appears to.
-
Interpretability
An honest account of what interpretability can do
Sparse autoencoders gave the field real traction. They did not give it the ability to certify a model as safe, and the distinction matters.
-
Alignment
Alignment is not one problem
Specification, robustness and assurance get collapsed into a single word, and the conflation quietly wrecks otherwise sensible arguments.