15 papers
The Capability Frontier: Benchmarks Miss 82% of Model Performance
Bradley Fowler, Ryan Smith, Daniel Thi Graviet +8
Existing benchmarks typically report accuracy for a single model on a single run. This systematically understates real-world LLM capabilities, particularly under heterogeneous data…
Are LLM Evaluators Really Narcissists? Sanity Checking Self-Preference Evaluations
Dani Roytburg, Matthew Bozoukov, Matthew Nguyen +3
Recent research has shown that large language models (LLMs) favor their own outputs when acting as judges, undermining the integrity of automated post-training and evaluation workf…
Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators
Dani Roytburg, Matthew Bozoukov, Matthew Nguyen +3
Large language models (LLMs) increasingly serve as automated evaluators, yet they suffer from "self-preference bias": a tendency to favor their own outputs over those of other mode…
Riemannian-Manifold Steering: Geometry-Aware Generative Autoencoders for Label-Free Steering
Narmeen Oozeer, Shivam Raval, Philip Quirke +4
Steering a language model - intervening on its internal activations to change downstream behaviour - has recently expanded beyond linear interpolation to nonlinear methods such as…
Make Mechanistic Interpretability Auditable: A Call to Develop Guidelines via Continuous Collaborative Reviewing
Michael Lan, Narmeen Fatimah Oozeer, Chaithanya Bandi +4
While mechanistic interpretability (MI) has produced important insights into neural network internals, the field has yet to establish a standardized system to audit experiments. As…
Geometry-Aware CLIP Retrieval via Local Cross-Modal Alignment and Steering
Nirmalendu Prakash, Narmeen Fatimah Oozeer, Xin Su +8
CLIP retrieval is typically framed as a pointwise similarity problem in a shared embedding space. While CLIP achieves strong global cross-modal alignment, many retrieval failures a…