3 papers
cs.LG2026
Do Models Read What They Write? Causal Registers in Scratchpad Reasoning
Benjamin Shih, John Winnicki, Eric Darve
A central hope behind process supervision is that models can expose intermediate variables that matter for their later behavior. For this to help with alignment, a scratchpad must…
cs.AI2026
Domain-Filtered Knowledge Graphs from Sparse Autoencoder Features
John Winnicki, Abeynaya Gnanasekaran, Eric Darve
Sparse autoencoders (SAEs) extract millions of interpretable features from a language model, but flat feature inventories aren't very useful on their own. Domain concepts get mixed…
physics.acc-ph2025
Coincident Learning for Beam-based RF Station Fault Identification Using Phase Information at the SLAC Linac Coherent Light Source
Jia Liang, William Colocho, Franz-Josef Decker +7
Anomalies in radio-frequency (RF) stations can result in unplanned downtime and performance degradation in linear accelerators such as SLAC's Linac Coherent Light Source (LCLS). De…