Showing cs.AIShow all
3 papers · 1 filter
cs.AI2025
Approximating Human Preferences Using a Multi-Judge Learned System
Eitán Sprejer, Fernando Avalos, Augusto Bernardi +3
Aligning LLM-based judges with human preferences is a significant challenge, as they are difficult to calibrate and often suffer from rubric sensitivity, bias, and instability. Ove…
cs.AI2025
Position: Require Frontier AI Labs To Release Small "Analog" Models
Shriyash Upadhyay, Chaithanya Bandi, Narmeen Oozeer +1
Recent proposals for regulating frontier AI models have sparked concerns about the cost of safety regulation, and most such regulations have been shelved due to the safety-innovati…
cs.AI2025
Activation Space Interventions Can Be Transferred Between Large Language Models
Narmeen Oozeer, Dhruv Nathawani, Nirmalendu Prakash +3
The study of representation universality in AI models reveals growing convergence across domains, modalities, and architectures. However, the practical applications of representati…