13 papers
CausalDS: Benchmarking Causal Reasoning in Data-Science Agents
Andrej Leban, Yuekai Sun
Large language models (LLMs) increasingly act as integrated data-science agents, combining abstract reasoning with advanced tool use. Yet the relevant benchmark landscape largely d…
From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning
Lingjing Kong, Xin Liu, Guangyi Chen +9
Post-training pipelines that combine supervised fine-tuning (SFT) with reinforcement learning (RL) have emerged as the key recipe for transforming large language models (LLMs) into…
LeanMarathon: Toward Reliable AI Co-Mathematicians through Long-Horizon Lean Autoformalization
Yuanhe Zhang, Yuekai Sun, Taiji Suzuki +2
Long-horizon autoformalization of research mathematics fails not only at hard lemmas, but at scale: statements drift, dependencies tangle, context decays, and local repairs corrupt…
A Latent Variable Framework for Scaling Laws in Large Language Models
Peiyao Cai, Chengyu Cui, Felipe Maia Polo +6
We propose a statistical framework built on latent variable modeling for scaling laws of large language models (LLMs). Our work is motivated by the rapid emergence of numerous new…
The Open Source Economic Index of AI Adoption and Capability
Seamus Somerstep, Aritra Guha, Divesh Srivastava +1
We work towards measuring both AI adoption and the capability of AI to perform discrete labor tasks across various occupations. To measure adoption, we develop an open-source econo…
Learning to Choose or Choosing to Learn: Best-of-N vs. Supervised Fine-Tuning for Bit String Generation
Seamus Somerstep, Vinod Raman, Unique Subedi +1
Using the bit string generation problem as a case study, we theoretically compare two standard methods for adapting large language models to new tasks. The first, referred to as su…