8 papers
Harnessing the Collective Intelligence of AI Agents in the Wild for New Discoveries
Federico Bianchi, Yongchan Kwon, Aneesh Pappu +1
Scientific discovery is often a collective process: researchers share partial results, inspect failed attempts, and build on each other's ideas over long time horizons. Recent AI s…
Automated Benchmark Auditing for AI Agents and Large Language Models
Junlin Wang, Federico Bianchi, Shang Zhu +4
Modern AI benchmarks operate at a complexity that outpaces traditional verification methods. Tasks authored by domain experts often contain implicit assumptions, incomplete environ…
Voice "Cloning" is Style Transfer
Kaitlyn Zhou, Federico Bianchi, Martijn Bartelds +3
Artificially generated speech is increasingly embedded in everyday life. Voice cloning in particular enables applications where identity preservation is important, such as completi…
What LLMs Think When You Don't Tell Them What to Think About?
Yongchan Kwon, James Zou
Characterizing the behavior of large language models (LLMs) across diverse settings is critical for reliable monitoring and AI safety. However, most existing analyses rely on topic…
DSGym: A Holistic Framework for Evaluating and Training Data Science Agents
Fan Nie, Junlin Wang, Harper Hua +6
Data science agents promise to accelerate discovery and insight-generation by turning data into executable analyses and findings. Yet existing data science benchmarks fall short du…
To Err Is Human: Systematic Quantification of Errors in Published AI Papers via LLM Analysis
Federico Bianchi, Yongchan Kwon, Zachary Izzo +2
How many mistakes do published AI papers contain? Peer-reviewed publications form the foundation upon which new research and knowledge are built. Errors that persist in the literat…