2 papers
cs.AI2025
Synthetic Error Injection Fails to Elicit Self-Correction In Language Models
David X. Wu, Shreyas Kapur, Anant Sahai +1
Reinforcement learning has become the dominant paradigm for eliciting reasoning and self-correction capabilities in large language models, but its computational expense motivates e…
cs.LG2024
Provable Weak-to-Strong Generalization via Benign Overfitting
David X. Wu, Anant Sahai
The classic teacher-student model in machine learning posits that a strong teacher supervises a weak student to improve the student's capabilities. We instead consider the inverted…