6 papers
Localizing RL-Induced Tool Use to a Single Crosscoder Feature
Andrii Shportko, Shubham Bhokare, Ahmed Zeyad A Alzahrani +3
Fine-tuning through RL reshapes the internal representations of language models to enable agentic behaviors such as tool use, yet the mechanistic basis of these changes remains poo…
Bridging Predictions and Interventions: An Integrated Framework for Automated Decision-Systems
Inioluwa Deborah Raji, Lydia T. Liu, Angela Zhou +27
Automated decision systems (ADS) leverage predictions about individual future outcomes to inform consequential decision-making in organizational settings. Across various settings -…
A Computational Method for Measuring "Open Codes" in Qualitative Analysis
John Chen, Alexandros Lotsos, Sihan Cheng +7
Qualitative analysis is critical to understanding human datasets in many social science disciplines. A central method in this process is inductive coding, where researchers identif…
Learning from the Right Rollouts: Data Attribution for PPO-based LLM Post-Training
Dong Shu, Denghui Zhang, Jessica Hullman
Traditional RL algorithms like Proximal Policy Optimization (PPO) typically train on the entire rollout buffer, operating under the assumption that all generated episodes provide a…
ComplLLM: Fine-tuning LLMs to Discover Complementary Signals for Decision-making
Ziyang Guo, Yifan Wu, Jason Hartline +2
Multi-agent decision pipelines can outperform single agent workflows when complementarity holds, i.e., different agents bring unique information to the table to inform a final deci…
How AI Responses Shape User Beliefs: The Effects of Information Detail and Confidence on Belief Strength and Stance
Zekun Wu, Mayank Jobanputra, Vera Demberg +2
The growing use of AI-generated responses in everyday tools raises concern about how subtle features such as supporting detail or tone of confidence may shape people's beliefs. To…