1 citations · 1 across the 3 of their papers we have counts for
3 papers
What Can You Do When You Have Zero Rewards During RL?
Jatin Prakash, Anirudh Buvanesh
Reinforcement learning (RL) with outcome-based rewards has proven effective for improving large language models (LLMs) on complex reasoning tasks. However, its success often depend…
On the Necessity of World Knowledge for Mitigating Missing Labels in Extreme Classification
Jatin Prakash, Anirudh Buvanesh, Bishal Santra +6
Extreme Classification (XC) aims to map a query to the most relevant documents from a very large document set. XC algorithms used in real-world applications learn this mapping from…
A Novel Data Augmentation Technique for Out-of-Distribution Sample Detection using Compounded Corruptions
Ramya S. Hebbalaguppe, Soumya Suvra Goshal, Jatin Prakash +2
Modern deep neural network models are known to erroneously classify out-of-distribution (OOD) test data into one of the in-distribution (ID) training classes with high confidence.…