Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
What LLMs Think When You Don't Tell Them What to Think About?
Yongchan Kwon, James Zou
Characterizing the behavior of large language models (LLMs) across diverse settings is critical for reliable monitoring and AI safety. However, most existing analyses rely on topic…
cs.AI2026
DSGym: A Holistic Framework for Evaluating and Training Data Science Agents
Fan Nie, Junlin Wang, Harper Hua +6
Data science agents promise to accelerate discovery and insight-generation by turning data into executable analyses and findings. Yet existing data science benchmarks fall short du…
cs.AI2025
To Err Is Human: Systematic Quantification of Errors in Published AI Papers via LLM Analysis
Federico Bianchi, Yongchan Kwon, Zachary Izzo +2
How many mistakes do published AI papers contain? Peer-reviewed publications form the foundation upon which new research and knowledge are built. Errors that persist in the literat…