2 papers
cs.AI2026
CORE: Contrastive Reflection Enables Rapid Improvements in Reasoning
Linas Nasvytis, Simon Jerome Han, Ben Prystawski +3
Language models can use verifiable rewards to improve at a wide variety of reasoning tasks. However, both parametric (e.g. RLVR) and non-parametric (e.g. prompt optimization) appro…
cs.LG2026
Addressing divergent representations from causal interventions on neural networks
Satchel Grant, Simon Jerome Han, Alexa R. Tartaglini +1
A common approach to mechanistic interpretability is to causally manipulate model representations via targeted interventions in order to understand what those representations encod…