2 papers
cs.CL2026
On the Limits of Steering Vectors for Preference-Aligned Generation
Melanie Subbiah, Zara Hall, Kathleen McKeown
Steering vectors have emerged as a promising approach to controlled text generation, offering interpretable, training-free mechanisms for shaping model outputs. However, their prac…
cs.LG2025
Guiding LLM Decision-Making with Fairness Reward Models
Zara Hall, Melanie Subbiah, Thomas P Zollo +2
Large language models are increasingly used to support high-stakes decisions, potentially influencing who is granted bail or receives a loan. Naive chain-of-thought sampling can im…