From the 1 of 8 linked papers with an AI index.
8 papers
OptimismBench: Forecasting Bias and the Alignment Effect in Language Model Judgment
Seonglae Cho, Adriano Koshiyama
The paper presents OptimismBench, a benchmark that measures directional optimism or pessimism in large language models' probability judgments by comparing paired success/failure fo…
Are Single-Token Sparse Autoencoder Features Causally Necessary? Layer-Depth and SAE-Family Effects
Seonglae Cho, Zekun Wu, Kleyton Da Costa +3
Sparse autoencoder (SAE) features are used to interpret and steer large language models, yet whether a feature's causal role is stable across SAE families remains untested. Single-…
Tool Calling is Linearly Readable and Steerable in Language Models
Zekun Wu, Ze Wang, Seonglae Cho +4
When a tool-calling agent picks the wrong tool, the failure is invisible until execution: the email gets sent, the meeting gets missed. As agents take on consequential actions, one…
Control Reinforcement Learning: Interpretable Token-Level Steering of LLMs via Sparse Autoencoder Features
Seonglae Cho, Zekun Wu, Adriano Koshiyama
Sparse autoencoders (SAEs) decompose language model activations into interpretable features, but existing methods reveal only which features activate, not which change model output…
CorrSteer: Generation-Time LLM Steering via Correlated Sparse Autoencoder Features
Seonglae Cho, Zekun Wu, Adriano Koshiyama
Sparse Autoencoders (SAEs) can extract interpretable features from large language models (LLMs) without supervision. However, their effectiveness in downstream steering tasks is li…
The Confidence Manifold: Geometric Structure of Correctness Representations in Language Models
Seonglae Cho, Zekun Wu, Kleyton Da Costa +1
When a language model asserts that "the capital of Australia is Sydney," does it know this is wrong? Models assert misconceptions with the same fluency as facts, so the question ca…