3 papers
cs.LG2026
Learning to Think from Multiple Thinkers
Nirmit Joshi, Roey Magen, Nathan Srebro +2
We study learning with Chain-of-Thought (CoT) supervision from multiple thinkers, all of whom provide correct but possibly systematically different solutions, e.g., step-by-step so…
cs.LG2025
Transformers are almost optimal metalearners for linear classification
Roey Magen, Gal Vardi
Transformers have demonstrated impressive in-context learning (ICL) capabilities, raising the question of whether they can serve as metalearners that adapt to new tasks using only…
cs.LG2025
Benign Overfitting in Single-Head Attention
Roey Magen, Shuning Shang, Zhiwei Xu +3
The phenomenon of benign overfitting, where a trained neural network perfectly fits noisy training data but still achieves near-optimal test performance, has been extensively studi…