2 citations · 2 across the 1 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
What Can You Do When You Have Zero Rewards During RL?
Jatin Prakash, Anirudh Buvanesh
Reinforcement learning (RL) with outcome-based rewards has proven effective for improving large language models (LLMs) on complex reasoning tasks. However, its success often depend…
cs.LG2024
On the Necessity of World Knowledge for Mitigating Missing Labels in Extreme Classification
Jatin Prakash, Anirudh Buvanesh, Bishal Santra +6
Extreme Classification (XC) aims to map a query to the most relevant documents from a very large document set. XC algorithms used in real-world applications learn this mapping from…