3 papers
cs.AI2026
Climbing the Ladder of Reasoning: What LLMs Can-and Still Can't-Solve after SFT?
Yiyou Sun, Georgia Zhou, Haoyue Bai +4
Recent supervised fine-tuning (SFT) approaches have significantly improved language models' performance on mathematical reasoning tasks, even when models are trained at a small sca…
cs.CL2025
OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization
Yiyou Sun, Shawn Hu, Georgia Zhou +4
Recent large-scale language models (LLMs) with long Chain-of-Thought reasoning-such as DeepSeek-R1-have achieved impressive results on Olympiad-level mathematics benchmarks. Howeve…
cs.AI2025
Efficient Automated Circuit Discovery in Transformers using Contextual Decomposition
Aliyah R. Hsu, Georgia Zhou, Yeshwanth Cherapanamjeri +4
Automated mechanistic interpretation research has attracted great interest due to its potential to scale explanations of neural network internals to large models. Existing automate…