17 papers
Measuring and mitigating overreliance to build human-compatible AI
Lujain Ibrahim, Katherine M. Collins, Sunnie S. Y. Kim +14
Large language models (LLMs) distinguish themselves from previous technologies by functioning as collaborative ``thought partners,'' capable of engaging more fluidly in natural lan…
Medical Model Synthesis Architectures: A Case Study
Katherine M. Collins, Marlene Berke, Ilia Sucholutsky +6
Medicine is rife with high-stakes uncertainty. Doctors routinely make clinical judgments and decisions that juggle many fundamental unknowns, like predictions about what might be c…
Improving the Efficiency of Language Agent Teams with Adaptive Task Graphs
Elizabeth Mieczkowski, Alexander Ku, Tiwalayo Eisape +5
Large language models (LLMs) are increasingly deployed in teams, yet existing coordination approaches often occupy two extremes. Highly structured methods rely on fixed roles, pipe…
Language Model Teams as Distributed Systems
Elizabeth Mieczkowski, Katherine M. Collins, Ilia Sucholutsky +2
Large language models (LLMs) are growing increasingly capable, prompting recent interest in LLM teams. Yet, despite increased deployment of LLM teams at scale, we lack a principled…
Data for Mathematical Copilots: Better Ways of Presenting Proofs for Machine Learning
Simon Frieder, Jonas Bayer, Sam Looi +13
The datasets and benchmarks commonly used to train and evaluate the mathematical capabilities of AI-based mathematical copilots (primarily large language models) exhibit several sh…
Documenting Deployment with Fabric: A Repository of Real-World AI Governance
Mackenzie Jorgensen, Kendall Brogle, Katherine M. Collins +10
Artificial intelligence (AI) is increasingly integrated into society, from financial services and traffic management to creative writing. Academic literature on the deployment of A…