activity
20242026
collaborators
Showing cs.CLShow all

6 papers · 1 filter

cs.CL2026

Expected Reward Prediction, with Applications to Model Routing

Kenan Hasanaliyev, Silas Alberti, Jenny Hamer +5

Reward models are a standard tool to score responses from LLMs. Reward models are built to rank responses to a fixed prompt sampled from a single model, for example to choose the b…

cs.CL2025

Steering off Course: Reliability Challenges in Steering Language Models

Patrick Queiroz Da Silva, Hari Sethuraman, Dheeraj Rajagopal +2

Steering methods for language models (LMs) have gained traction as lightweight alternatives to fine-tuning, enabling targeted modifications to model activations. However, prior stu…

cs.CL2025

AutoMix: Automatically Mixing Language Models

Pranjal Aggarwal, Aman Madaan, Ankit Anand +10

Large language models (LLMs) are now available from cloud API providers in various sizes and configurations. While this diversity offers a broad spectrum of choices, effectively le…

cs.CL2024

Scalable Influence and Fact Tracing for Large Language Model Pretraining

Tyler A. Chang, Dheeraj Rajagopal, Tolga Bolukbasi +2

Training data attribution (TDA) methods aim to attribute model outputs back to specific training examples, and the application of these methods to large language model (LLM) output…

cs.CL2024

How Far Can We Extract Diverse Perspectives from Large Language Models?

Shirley Anugrah Hayati, Minhwa Lee, Dheeraj Rajagopal +1

Collecting diverse human opinions is costly and challenging. This leads to a recent trend in exploiting large language models (LLMs) for generating diverse data for potential scala…

cs.CL2024

Confidence Calibration and Rationalization for LLMs via Multi-Agent Deliberation

Ruixin Yang, Dheeraj Rajagopal, Shirley Anugrah Hayati +2

Uncertainty estimation is a significant issue for current large language models (LLMs) that are generally poorly calibrated and over-confident, especially with reinforcement learni…