4 papers
Boule or Baguette? A Study on Task Topology, Length Generalization, and the Benefit of Reasoning Traces
William L. Tong, Ege Cakar, Cengiz Pehlevan
Recent years have witnessed meteoric progress in reasoning models: neural networks that generate intermediate reasoning traces (RTs) before producing a final output. Despite the ra…
No Free Lunch From Random Feature Ensembles: Scaling Laws and Near-Optimality Conditions
Benjamin S. Ruben, William L. Tong, Hamza Tahir Chaudhry +1
Given a fixed budget for total model size, one must choose between training a single large model or combining the predictions of multiple smaller models. We investigate this trade-…
Learning richness modulates equality reasoning in neural networks
William L. Tong, Cengiz Pehlevan
Equality reasoning is ubiquitous and purely abstract: sameness or difference may be evaluated no matter the nature of the underlying objects. As a result, same-different (SD) tasks…
MLPs Learn In-Context on Regression and Classification Tasks
William L. Tong, Cengiz Pehlevan
In-context learning (ICL), the remarkable ability to solve a task from only input exemplars, is often assumed to be a unique hallmark of Transformer models. By examining commonly e…