3 papers
cs.AI2026
Improving Evaluation Realism with Inference-Time Compute and Deployment Scaffolds
Axel Ahlqvist, Richard Guan, Juan-Pablo Rivera +6
A core obstacle to alignment evaluation is evaluation awareness: capable models can tell when they are being tested rather than deployed, weakening the conclusions a safety evaluat…
cs.AI2025
Internal states before wait modulate reasoning patterns
Dmitrii Troitskii, Koyena Pal, Chris Wendler +2
Prior work has shown that a significant driver of performance in reasoning models is their ability to reason and self-correct. A distinctive marker in these reasoning traces is the…
cs.LG2024
NNsight and NDIF: Democratizing Access to Open-Weight Foundation Model Internals
Jaden Fiotto-Kaufman, Alexander R. Loftus, Eric Todd +17
We introduce NNsight and NDIF, technologies that work in tandem to enable scientific study of the representations and computations learned by very large neural networks. NNsight is…