Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
CausalDS: Benchmarking Causal Reasoning in Data-Science Agents
Andrej Leban, Yuekai Sun
Large language models (LLMs) increasingly act as integrated data-science agents, combining abstract reasoning with advanced tool use. Yet the relevant benchmark landscape largely d…
cs.AI2026
LeanMarathon: Toward Reliable AI Co-Mathematicians through Long-Horizon Lean Autoformalization
Yuanhe Zhang, Yuekai Sun, Taiji Suzuki +2
Long-horizon autoformalization of research mathematics fails not only at hard lemmas, but at scale: statements drift, dependencies tangle, context decays, and local repairs corrupt…