190 citations · 216 across the 5 of their papers we have counts for
4 papers
Stepwise Penalization for Length-Efficient Chain-of-Thought Reasoning
Xintong Li, Sha Li, Rongmei Lin +10
Large reasoning models improve with more test-time computation, but often overthink, producing unnecessarily long chains-of-thought that raise cost without improving accuracy. Prio…
Exploring the MIT Mathematics and EECS Curriculum Using Large Language Models
Sarah J. Zhang, Samuel Florin, Ariel N. Lee +12
We curate a comprehensive dataset of 4,550 questions and solutions from problem sets, midterm exams, and final exams across all MIT Mathematics and Electrical Engineering and Compu…
From Human Days to Machine Seconds: Automatically Answering and Generating Machine Learning Final Exams
Iddo Drori, Sarah J. Zhang, Reece Shuttleworth +13
A final exam in machine learning at a top institution such as MIT, Harvard, or Cornell typically takes faculty days to write, and students hours to solve. We demonstrate that large…
A Neural Network Solves, Explains, and Generates University Math Problems by Program Synthesis and Few-Shot Learning at Human Level
Iddo Drori, Sarah Zhang, Reece Shuttleworth +15
We demonstrate that a neural network pre-trained on text and fine-tuned on code solves mathematics course problems, explains solutions, and generates new questions at a human level…