11 papers
On Incentivized Exploration beyond Bayesianism and Full-Information
Dimitar Chakarov, Lee Cohen, Nathan Srebro
We extend Incentive Compatible Exploration beyond the Bayesian full-information setting of Kremer et al. [2014]. We consider agents that may possess external information unknown to…
Recursive Models for Long-Horizon Reasoning
Chenxiao Yang, Nathan Srebro, Zhiyuan Li
Modern language models reason within bounded context, an inherent constraint that poses a fundamental barrier to long-horizon reasoning. We identify recursion as a core principle f…
Learning through Internalization
Nikolaos Tsilivis, Nirmit Joshi, Marko Medvedev +2
We study internalization processes, by which neural-network-based systems absorb an explicit computational procedure into their own weights, and how they facilitate learning. We in…
Tight Sample Complexity of Transformers
Chenxiao Yang, Nathan Srebro, Zhiyuan Li
We tightly characterize the VC dimension of depth- Transformers with a total of parameters, mapping an input sequence of length to a single output, establishing an upper…
Learning to Think from Multiple Thinkers
Nirmit Joshi, Roey Magen, Nathan Srebro +2
We study learning with Chain-of-Thought (CoT) supervision from multiple thinkers, all of whom provide correct but possibly systematically different solutions, e.g., step-by-step so…
Learning to Answer from Correct Demonstrations
Nirmit Joshi, Gene Li, Siddharth Bhandari +3
We study the problem of learning to generate an answer (or completion) to a question (or prompt), where there could be multiple correct answers, any one of which is acceptable at t…