2 citations · 3 across the 13 of their papers we have counts for
4 papers · 1 filter
LLM Detection as an Intervention: Downstream Impact under Strategic User Behavior
Meena Jagadeesan, Tatsunori Hashimoto, Jon Kleinberg
As LLM adoption becomes more widespread, there is a growing interest in detecting LLM-generated content, for example through LLM detection tools and through heuristics based on lan…
Metric Match: A Subset Selection Approach to Evaluating LLM Judge Reliability
Alyssa Unell, Natalie Dullerud, Naomi Boneh +4
LLM judges are used to reduce the need for costly human labor in evaluating open-ended text generation. However, the reliability of these judges depends critically on their alignme…
Deployment-Centered Evaluation: Predicting Query-Level Rejection Risk in a Clinical LLM System
Alyssa Unell, Miguel Fuentes, Brenna Li +4
Large language models (LLMs) are increasingly integrated into clinical systems, making it essential to evaluate the real-world utility of these systems. However, static benchmarks…
Power and Limitations of Aggregation in Compound AI Systems
Nivasini Ananthakrishnan, Meena Jagadeesan
When designing compound AI systems, a common approach is to query multiple copies of the same model and aggregate the responses to produce a synthesized output. Given the homogenei…