7 papers
Spectral Superposition: A Theory of Feature Geometry
Georgi Ivanov, Narmeen Oozeer, Shivam Raval +3
Neural networks represent more features than they have dimensions via superposition, forcing features to share representational space. Current methods decompose activations into sp…
Approximating Human Preferences Using a Multi-Judge Learned System
Eitán Sprejer, Fernando Avalos, Augusto Bernardi +3
Aligning LLM-based judges with human preferences is a significant challenge, as they are difficult to calibrate and often suffer from rubric sensitivity, bias, and instability. Ove…
Position: Require Frontier AI Labs To Release Small "Analog" Models
Shriyash Upadhyay, Chaithanya Bandi, Narmeen Oozeer +1
Recent proposals for regulating frontier AI models have sparked concerns about the cost of safety regulation, and most such regulations have been shelved due to the safety-innovati…
Distribution-Aware Feature Selection for SAEs
Narmeen Oozeer, Nirmalendu Prakash, Michael Lan +2
Sparse autoencoders (SAEs) decompose neural activations into interpretable features. A widely adopted variant, the TopK SAE, reconstructs each token from its K most active latents.…
Beyond Monoliths: Expert Orchestration for More Capable, Democratic, and Safe Language Models
Philip Quirke, Narmeen Oozeer, Chaithanya Bandi +8
This position paper argues that the prevailing trajectory toward ever larger, more expensive generalist foundation models controlled by a handful of companies limits innovation and…
Activation Space Interventions Can Be Transferred Between Large Language Models
Narmeen Oozeer, Dhruv Nathawani, Nirmalendu Prakash +3
The study of representation universality in AI models reveals growing convergence across domains, modalities, and architectures. However, the practical applications of representati…