Publications (13)
Instruct-SkillMix: A Powerful Pipeline for LLM Instruction Tuning
Simran Kaur, Simon Park, Anirudh Goyal +1
We introduce Instruct-SkillMix, an automated approach for creating diverse, high quality SFT data for instruction-following. The pipeline involves two stages, each leveraging an ex…
Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision
Yinghui He, Simran Kaur, Adithya Bhaskar +7
Current post-training methods in verifiable settings fall into two categories. Reinforcement learning (RLVR) relies on binary rewards, which are broadly applicable and powerful, bu…
Skill-Mix: a Flexible and Expandable Family of Evaluations for AI models
Dingli Yu, Simran Kaur, Arushi Gupta +3
With LLMs shifting their role from statistical modeling of language to serving as general-purpose AI agents, how should LLM evaluations change? Arguably, a key ability of an AI age…
Gradient Descent on Neural Networks Typically Occurs at the Edge of Stability
Jeremy M. Cohen, Simran Kaur, Yuanzhi Li +2
We empirically demonstrate that full-batch gradient descent on neural network training objectives typically operates in a regime we call the Edge of Stability. In this regime, the…
Rethinking On-Policy Self-Distillation for Thinking Models
Simran Kaur, Narutatsu Ri, Yinghui He +2
Self-distillation is a promising recipe for self-improvement in language models. In this setting, a model can serve as its own teacher when given privileged information, such as a…
A Joint Search for the Electromagnetic Counterpart to the Gravitational-Wave Binary Black-Hole Merger Candidate S250328ae with the Dark Energy Camera and the Prime Focus Spectrograph
Haibin Zhang, Mitsuru Kokubo, Sean MacBride +20
The first detection of an optical counterpart to a gravitational wave signal revealed that collaborative efforts between instruments with different specializations provide a unique…