10 papers
Societal Alignment Frameworks Can Improve LLM Alignment
Karolina StaÅczak, Nicholas Meade, Mehar Bhatia +14
Recent progress in large language models (LLMs) has focused on producing responses that meet human expectations and align with shared values - a process coined alignment. However,…
Confidence Calibration in Vision-Language-Action Models
Thomas P Zollo, Richard Zemel
Trustworthy robot behavior requires not only high levels of task success but also that the robot can reliably quantify how likely it is to succeed. To this end, we present a first-…
Let the Experts Speak: Improving Survival Prediction & Calibration via Mixture-of-Experts Heads
Todd Morrill, Aahlad Puli, Murad Megjhani +2
Deep mixture-of-experts models have attracted a lot of attention for survival analysis problems, particularly for their ability to cluster similar patients together. In practice, g…
Guiding LLM Decision-Making with Fairness Reward Models
Zara Hall, Melanie Subbiah, Thomas P Zollo +2
Large language models are increasingly used to support high-stakes decisions, potentially influencing who is granted bail or receives a loan. Naive chain-of-thought sampling can im…
Replay Can Provably Increase Forgetting
Yasaman Mahdaviyeh, James Lucas, Mengye Ren +3
Continual learning seeks to enable machine learning systems to solve an increasing corpus of tasks sequentially. A critical challenge for continual learning is forgetting, where th…
Towards Safety Reasoning in LLMs: AI-agentic Deliberation for Policy-embedded CoT Data Creation
Tharindu Kumarage, Ninareh Mehrabi, Anil Ramakrishna +6
Safety reasoning is a recent paradigm where LLMs reason over safety policies before generating responses, thereby mitigating limitations in existing safety measures such as over-re…