11 citations · 11 across the 4 of their papers we have counts for
3 papers · 1 filter
UNDO: Understanding Distillation as Optimization
Kushal Jain, Piyushi Goyal, Kumar Shridhar
Knowledge distillation has emerged as an effective strategy for compressing large language models' (LLMs) knowledge into smaller, more efficient student models. However, standard o…
First-Step Advantage: Importance of Starting Right in Multi-Step Math Reasoning
Kushal Jain, Moritz Miller, Niket Tandon +1
Language models can solve complex reasoning tasks better by learning to generate rationales for their predictions. Often these models know how to solve a task but their auto-regres…
Indic-Transformers: An Analysis of Transformer Language Models for Indian Languages
Kushal Jain, Adwait Deshpande, Kumar Shridhar +2
Language models based on the Transformer architecture have achieved state-of-the-art performance on a wide range of NLP tasks such as text classification, question-answering, and t…