6 papers
Do Models Read What They Write? Causal Registers in Scratchpad Reasoning
Benjamin Shih, John Winnicki, Eric Darve
A central hope behind process supervision is that models can expose intermediate variables that matter for their later behavior. For this to help with alignment, a scratchpad must…
Domain-Filtered Knowledge Graphs from Sparse Autoencoder Features
John Winnicki, Abeynaya Gnanasekaran, Eric Darve
Sparse autoencoders (SAEs) extract millions of interpretable features from a language model, but flat feature inventories aren't very useful on their own. Domain concepts get mixed…
Flash-SD-KDE: Accelerating SD-KDE with Tensor Cores
Elliot L. Epstein, Rajat Vadiraj Dwaraknath, John Winnicki
Score-debiased kernel density estimation (SD-KDE) achieves improved asymptotic convergence rates over classical KDE, but its use of an empirical score has made it significantly slo…
Allocate Marginal Reviews to Borderline Papers Using LLM Comparative Ranking
Elliot L. Epstein, Rajat Dwaraknath, John Winnicki +1
This paper argues that large ML conferences should allocate marginal review capacity primarily to papers near the acceptance boundary, rather than spreading extra reviews via rando…
LLMs are Overconfident: Evaluating Confidence Interval Calibration with FermiEval
Elliot L. Epstein, John Winnicki, Thanawat Sornwanee +1
Large language models (LLMs) excel at numerical estimation but struggle to correctly quantify uncertainty. We study how well LLMs construct confidence intervals around their own an…
SD-KDE: Score-Debiased Kernel Density Estimation
Elliot L. Epstein, Rajat Dwaraknath, Thanawat Sornwanee +2
We propose a novel method for density estimation that leverages an estimated score function to debias kernel density estimation (SD-KDE). In our approach, each data point is adjust…