1 citations · 1 across the 3 of their papers we have counts for
3 papers · 1 filter
Demystifying Reinforcement Learning Post-Training of Language Models
Donovan Clay, Saket Gollapudi, Sankar Harilal +4
Reinforcement learning (RL) post-training has emerged as a powerful framework for enhancing the capabilities of large language models (LLMs), enabling impressive reasoning, math, a…
Understanding the Gain from Data Filtering in Multimodal Contrastive Learning
Divyansh Pareek, Sewoong Oh, Simon S. Du
The success of modern multimodal representation learning relies on internet-scale datasets. Due to the low quality of a large fraction of raw web data, data curation has become a c…
Understanding the Gains from Repeated Self-Distillation
Divyansh Pareek, Simon S. Du, Sewoong Oh
Self-Distillation is a special type of knowledge distillation where the student model has the same architecture as the teacher model. Despite using the same architecture and the sa…