3 papers
cs.LG2026
Theoretical Limits of Language Model Alignment
Lucas Monteiro Paes, Natalie Mackraz, Barry-John Theobald +1
Language model (LM) alignment improves model outputs to reflect human preferences while preserving the capabilities of the base model. The most common alignment approaches are (i)…
cs.LG2026
DSO: Direct Steering Optimization for Bias Mitigation
Lucas Monteiro Paes, Nivedha Sivakumar, Yinong Oliver Wang +4
Generative models are often deployed to make decisions on behalf of users, such as vision-language models (VLMs) identifying which person in a room is a doctor to help visually imp…
cs.CL2025
ICX360: In-Context eXplainability 360 Toolkit
Dennis Wei, Ronny Luss, Xiaomeng Hu +6
Large Language Models (LLMs) have become ubiquitous in everyday life and are entering higher-stakes applications ranging from summarizing meeting transcripts to answering doctors'…