6 papers · 1 filter
Expected Reward Prediction, with Applications to Model Routing
Kenan Hasanaliyev, Silas Alberti, Jenny Hamer +5
Reward models are a standard tool to score responses from LLMs. Reward models are built to rank responses to a fixed prompt sampled from a single model, for example to choose the b…
Steering off Course: Reliability Challenges in Steering Language Models
Patrick Queiroz Da Silva, Hari Sethuraman, Dheeraj Rajagopal +2
Steering methods for language models (LMs) have gained traction as lightweight alternatives to fine-tuning, enabling targeted modifications to model activations. However, prior stu…
AutoMix: Automatically Mixing Language Models
Pranjal Aggarwal, Aman Madaan, Ankit Anand +10
Large language models (LLMs) are now available from cloud API providers in various sizes and configurations. While this diversity offers a broad spectrum of choices, effectively le…
Scalable Influence and Fact Tracing for Large Language Model Pretraining
Tyler A. Chang, Dheeraj Rajagopal, Tolga Bolukbasi +2
Training data attribution (TDA) methods aim to attribute model outputs back to specific training examples, and the application of these methods to large language model (LLM) output…
How Far Can We Extract Diverse Perspectives from Large Language Models?
Shirley Anugrah Hayati, Minhwa Lee, Dheeraj Rajagopal +1
Collecting diverse human opinions is costly and challenging. This leads to a recent trend in exploiting large language models (LLMs) for generating diverse data for potential scala…
Confidence Calibration and Rationalization for LLMs via Multi-Agent Deliberation
Ruixin Yang, Dheeraj Rajagopal, Shirley Anugrah Hayati +2
Uncertainty estimation is a significant issue for current large language models (LLMs) that are generally poorly calibrated and over-confident, especially with reinforcement learni…