Circuit learning at scale: what the latest interpretability tools actually deliver
Sparse autoencoders, cross-layer transcoders, and condensation training are making circuit discovery practical on frontier models. A survey of what works, what doesn't, and what's next.
Mechanistic interpretability has a scale problem. The core idea is straightforward enough: find the internal components responsible for a behaviour, trace how information flows between them, and verify the account by intervening. The difficulty is that frontier models have billions of parameters, and the methods that worked on GPT-2-class networks do not survive the jump.
A cluster of recent papers addresses this gap from several angles at once. CircuitLasso brings sparse linear regression to circuit discovery. Structured sparse autoencoders force cross-modal consistency in vision-language models. Circuit Condensation trains models to concentrate their behaviours into smaller, inspectable subgraphs. And CLT-Forge provides the infrastructure to run all of this at production scale. Taken together, they represent a genuine step forward in what the field can claim about the internals of models people actually deploy.
Why raw neurons are not enough
The fundamental obstacle to circuit discovery in large models is polysemanticity. Individual neurons respond to unrelated concepts. A single neuron in a language model might activate for both the token "bank" in a financial context and "bank" in a riverbank context, along with a dozen other patterns. Tracing a circuit through polysemantic neurons produces explanations that do not cohere, because each node in the circuit means several things at once.
Sparse autoencoders (SAEs) offered a partial solution. By decomposing activations into a much larger set of sparse, more interpretable features, SAEs produce directions in activation space that correspond to single concepts. Anthropic's "Scaling Monosemanticity" work demonstrated that this approach recovers interpretable features from production-scale models, including features related to safety-relevant behaviours [Templeton et al., 2024].
But SAEs introduced new problems. The feature set is enormous, making circuit search combinatorially expensive. Features in vision-language models are not consistent across modalities, so a feature that means one thing in text may fragment in the visual domain. And post-hoc decomposition, by definition, happens after training, meaning the features are not aligned with downstream objectives.
CircuitLasso: scalable circuit discovery
CircuitLasso, from Yin et al. [2026], attacks the computational cost directly. Previous intervention-based circuit learning methods required running the model forward and backward for every candidate edge in the circuit graph. With millions of SAE features, this becomes prohibitive.
The key insight is that circuit learning can be reframed as sparse linear regression. Given a target behaviour and a set of SAE features, CircuitLasso finds the minimal subset of features that causally contribute to that behaviour using regularised regression rather than exhaustive intervention. The structural accuracy matches state-of-the-art intervention-based methods, but at a fraction of the computational cost.
The practical result is that researchers can now trace how human-interpretable semantic features propagate through a model's layers and influence its predictions. On a domain-generalisation task, insights from the learned circuits were leveraged to achieve comparable performance at substantially lower cost [Yin et al., 2026]. That is a useful property: the circuit is not just an explanation, it is a tool for prediction.
Structured SAEs for vision-language models
A separate problem arises in multimodal settings. Standard SAEs trained on vision-language models like Qwen2.5-VL-7B produce features that are interpretable in text but fragmented in the visual domain. A feature that captures "dog" in text might activate for disjoint regions of an image, or fail to localise the concept at all.
Liao, Yang and Wei [2026] propose Structured Sparse Autoencoders (S²AE) to address this. The method groups image patches based on Transformer attention similarity and spatial proximity, then applies structured sparsity during SAE training. Exclusive sparsity disentangles features between groups; group sparsity enforces consistency within groups.
Evaluated on Qwen2.5-VL-7B-Instruct, S²AE achieves 6.06% improvement in semantic alignment (mIoU) and a 60.81-point improvement in representational efficiency (lower L0 norm) while maintaining reconstruction fidelity above 99% explained variance. Cross-modal analysis shows a 3.08% gain in semantic consistency and a 2.37% gain in monosemanticity scores across both modalities [Liao et al., 2026].
These are not headline numbers. They are incremental improvements. But the direction matters: features that are consistent across modalities are features you can actually trace through a model's reasoning, rather than features that happen to be interpretable in one narrow context.
Circuit Condensation: making circuits smaller
Even with efficient discovery methods, circuits in large models remain unwieldy. Hundreds of edges make them hard to inspect, compare, or verify. Senthil Kumar [2026] introduces Circuit Condensation, a post-training technique that concentrates model behaviours into smaller causal subgraphs.
The method works iteratively. In each round, low-attribution edges are pruned from the circuit and a low-rank adapter is trained to match the original model's behaviour using only the remaining edges. If task performance and general capability survive the pruning, the condensed circuit is kept. Across four behaviours and eight models, condensed circuits are smaller than the strongest frozen baseline in 30 of 32 settings, by 8.1x on average and up to 316x [Senthil Kumar, 2026].
A key finding is that weight updates, not search, drive the reduction. Repeating the search without weight updates produces larger circuits in 29 of 32 settings. This suggests the models themselves can be trained to concentrate their computations into inspectable subgraphs, which is a more tractable engineering problem than trying to find ever-better search algorithms over a fixed model.
On indirect object identification, condensation isolates 24 heads, 17 of them with documented roles, against 61 heads and 36 undocumented ones for the matched frozen circuit. That is a sufficient sub-circuit of the published mechanism rather than a full reconstruction, which is arguably more useful: it tells you which components are load-bearing.
Cross-layer transcoders and the infrastructure problem
Scaling these methods requires infrastructure that did not exist a year ago. Cross-layer transcoders (CLTs) improve on standard SAEs by sharing features across layers while preserving layer-specific decoding, yielding more compact representations. But training and analysing CLTs at scale is difficult.
CLT-Forge [Draye et al., 2026] provides an open-source library for end-to-end training and interpretability of CLTs. It integrates distributed training with model sharding and compressed activation caching, a unified automated interpretability pipeline for feature analysis, attribution graph computation, and a visualisation interface. This is plumbing work, but plumbing work that determines whether the methods described above can be applied to models beyond the ones the original authors tested on.
SpIn-ViT: interpretable by design
The approaches described so far are post-hoc: they decompose representations after the model has been trained. SpIn-ViT [Lee and Padalkar, 2026] takes a different path for vision transformers, jointly training a ViT and a modified SAE end-to-end so that sparse, patch-level representations are directly aligned with the classification objective.
Compared with state-of-the-art post-hoc SAE methods, SpIn-ViT achieves 8.84% higher average classification accuracy, an AI-based interpretability score nearly four times as high, and a human evaluation score more than twice as high [Lee and Padalkar, 2026]. The method also extracts interpretable rule-sets from SAE neurons to create neurosymbolic models that achieve 5.97% higher accuracy with 58.8% smaller rule-sets.
This is the most promising direction for safety applications. If a model is trained from the start to have interpretable internal structure, the interpretability is not an afterthought bolted on by researchers. It is a design constraint. The trade-off is that it requires access to the training process, which limits it to organisations training their own models.
ObserverBench: measuring whether interpretability helps
A persistent question in the field is whether better internal estimates actually lead to better decisions. ObserverBench [Erramilli, 2026] tests this directly by evaluating whether internal estimators are adequate for the interventions they guide.
The findings are sobering. On circuit-intervention tasks in GPT-2-small and Qwen2.5-7B, pairwise observers predict unseen effects more accurately without always choosing better actions. Observers trained on action loss choose lower-loss actions, but the relationship between estimation accuracy and deployment performance is not straightforward.
Across Qwen2.5-7B, Gemma-2-9B-it, and Qwen3.5-9B, AUROC can rank monitors differently from deployment loss, and the best information source changes across models. Sparse SAE readouts also trail their layer-matched dense controls on certain panels, under disclosed activation-density or checkpoint mismatches [Erramilli, 2026].
This matters because interpretability methods are increasingly proposed as components of safety evaluation. If a method predicts well but acts poorly, or if its predictions do not transfer across models, then its value in a safety case is limited. ObserverBench provides the fixed task contracts and baselines needed to evaluate this rigorously.
What this means for safety and governance
The practical implications of these results are nuanced. On the positive side:
Circuit discovery is now feasible on frontier models. The combination of CircuitLasso for efficient search, Circuit Condensation for pruning, and CLT-Forge for infrastructure means the field can move beyond GPT-2-scale demonstrations. The S²AE work extends this to multimodal models, which are the ones being deployed at scale.
The ObserverBench results are a necessary corrective. Interpretability methods need to be evaluated by the actions they enable, not just by the accuracy of their internal estimates. The fact that AUROC does not reliably predict deployment loss should give pause to anyone proposing to use sparse SAE readouts as safety monitors without task-specific validation.
The EU AI Act's transparency requirements and the UK AI Security Institute's evaluation methodology both assume that model internals can be inspected. These papers provide evidence that the assumption is becoming less unreasonable, but they also show that the inspection methods have real limitations that must be disclosed.
A practical note on coverage. Every method described here recovers some circuits or features, not all of them. The uninterpreted residual in a model's computation is not established to be safety-irrelevant. This is the same caveat from our earlier piece on interpretability, and it remains load-bearing. These tools are getting better. They are not yet sufficient for safety certification.
What comes next
Three developments to watch:
First, whether Circuit Condensation generalises to safety-relevant behaviours. The current results cover indirect object identification and a small set of other tasks. Jailbreaking, deceptive alignment, and reward hacking are the behaviours that matter for safety, and it is not obvious that they concentrate into small circuits in the same way.
Second, whether SpIn-ViT-style approaches can be extended to language models. Joint training for interpretability is compelling in vision transformers, where the classification objective is clear. Language models have more diffuse objectives, and it is not clear how to impose the same kind of structured sparsity constraint.
Third, whether the infrastructure tools (CLT-Forge, CircuitLasso) get adoption outside the groups that built them. The field has a history of methods that work in the authors' hands but do not reproduce elsewhere. Open-source code helps, but reproducibility depends on documentation, stable APIs, and community investment.
The trend line is positive. Two years ago, mechanistic interpretability on frontier models was aspirational. Now it is difficult but feasible. The gap between "feasible" and "reliable enough for safety cases" remains wide, and honesty about that gap is more useful than either dismissiveness or hype.
References
Yin, N., Wei, D., Gao, T., Dhurandhar, A., Natesan Ramamurthy, K. and Yu, Y. (2026). "Scalable Circuit Learning for Interpreting Large Language Models." arXiv:2606.16939. https://arxiv.org/abs/2606.16939
Senthil Kumar, S.A. (2026). "Circuit Condensation: Post-Training that Concentrates a Behavior's Causal Circuit." arXiv:2608.27254. https://arxiv.org/abs/2608.27254
Lee, P.H. and Padalkar, P. (2026). "SpIn-ViT: Designing a Sparsity-Induced Vision Transformer That Is Mechanistically Interpretable." arXiv:2608.14922. https://arxiv.org/abs/2608.14922
Erramilli, V. (2026). "ObserverBench: Testing Mechanistic Estimates for Intervention and Control." arXiv:2609.03026. https://arxiv.org/abs/2609.03026
Draye, F. et al. (2026). "CLT-Forge: A Scalable Library for Cross-Layer Transcoders and Attribution Graphs." arXiv:2603.21014. https://arxiv.org/abs/2603.21014
Templeton, A. et al. (2024). "Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet." Anthropic Research. https://www.anthropic.com/research/mapping-mind-language-model
Liao, W., Yang, Y. and Wei, Y. (2026). "When Structured Sparse Autoencoders Learn Consistent Concepts Across Modalities." arXiv:2607.08605. https://arxiv.org/abs/2607.08605