7 papers
LLM Detection as an Intervention: Downstream Impact under Strategic User Behavior
Meena Jagadeesan, Tatsunori Hashimoto, Jon Kleinberg
As LLM adoption becomes more widespread, there is a growing interest in detecting LLM-generated content, for example through LLM detection tools and through heuristics based on lan…
Power and Limitations of Aggregation in Compound AI Systems
Nivasini Ananthakrishnan, Meena Jagadeesan
When designing compound AI systems, a common approach is to query multiple copies of the same model and aggregate the responses to produce a synthesized output. Given the homogenei…
Strategic Feature Selection
Jivat Neet Kaur, Pratik Patil, Divya Shanmugam +6
When algorithmic predictors inform resource allocation in high-stakes domains such as healthcare, these predictors must account for strategic manipulation of input features. The ty…
Metric Match: A Subset Selection Approach to Evaluating LLM Judge Reliability
Alyssa Unell, Natalie Dullerud, Naomi Boneh +4
LLM judges are used to reduce the need for costly human labor in evaluating open-ended text generation. However, the reliability of these judges depends critically on their alignme…
Deployment-Centered Evaluation: Predicting Query-Level Rejection Risk in a Clinical LLM System
Alyssa Unell, Miguel Fuentes, Brenna Li +4
Large language models (LLMs) are increasingly integrated into clinical systems, making it essential to evaluate the real-world utility of these systems. However, static benchmarks…
Breaking Algorithmic Collusion in Human-AI Ecosystems
Natalie Collina, Eshwar Ram Arunachaleswaran, Meena Jagadeesan
AI agents are increasingly deployed in ecosystems where they repeatedly interact not only with each other but also with humans. In this work, we study these human-AI ecosystems fro…