6 papers
Position: AI Evaluations Should be Grounded on a Theory of Capability
Nathanael Jo, Ashia Wilson
Evaluations of generative models are now ubiquitous, and their outcomes critically shape public and scientific expectations of AI's capabilities. Yet skepticism about their reliabi…
Evaluation without Generation: Non-Generative Assessment of Harmful Model Specialization with Applications to CSAM
Vinith M. Suriyakumar, Ayush Sekhari, Lena Stempfle +5
Auditing the fine-tunes of open-weight generative models for harmful specialization has become a new governance challenge for model hosting platforms. The standard toolkit, generat…
Online Reasoning Calibration: Test-Time Training Enables Generalizable Conformal LLM Reasoning
Cai Zhou, Zekai Wang, Menghua Wu +6
While test-time scaling has enabled large language models to solve highly difficult tasks, state-of-the-art results come at exorbitant compute costs. These inefficiencies can be at…
Efficient and accurate steering of Large Language Models through attention-guided feature learning
Parmida Davarmanesh, Ashia Wilson, Adityanarayanan Radhakrishnan
Steering, or direct manipulation of internal activations to guide LLM responses toward specific semantic concepts, is emerging as a promising avenue for both understanding how sema…
UCD: Unlearning in LLMs via Contrastive Decoding
Vinith M. Suriyakumar, Ayush Sekhari, Ashia Wilson
Machine unlearning aims to remove specific information, e.g. sensitive or undesirable content, from large language models (LLMs) while preserving overall performance. We propose an…
Layered Unlearning for Adversarial Relearning
Timothy Qian, Vinith Suriyakumar, Ashia Wilson +1
Our goal is to understand how post-training methods, such as fine-tuning, alignment, and unlearning, modify language model behavior and representations. We are particularly interes…