6 citations · 16 across the 15 of their papers we have counts for
9 papers · 1 filter
Predictable GRPO: A Closed-Form Model of Training Dynamics
Rajat Ghosh, Datta Nimmaturi, Aryan Singhal +4
We develop a first-principles reduced-order model of these dynamics. Under a single mean-field assumption that summarizes the policy by its expected reward, we reduce the GRPO upda…
Action Shapley: A Training Data Selection Metric for World Model in Reinforcement Learning
Rajat Ghosh, Debojyoti Dutta
Numerous offline and model-based reinforcement learning systems incorporate world models to emulate the inherent environments. A world model is particularly important in scenarios…
Go-UT-Bench: A Fine-Tuning Dataset for LLM-Based Unit Test Generation in Go
Yashshi Pipalani, Hritik Raj, Rajat Ghosh +2
Training data imbalance poses a major challenge for code LLMs. Most available data heavily over represents raw opensource code while underrepresenting broader software engineering…
A Multi-Agent Framework for Stateful Inference-Time Search
Arshika Lalan, Rajat Ghosh, Aditya Kolsur +1
Recent work explores agentic inference-time techniques to perform structured, multi-step reasoning. However, stateless inference often struggles on multi-step tasks due to the abse…
BAR Conjecture: the Feasibility of Inference Budget-Constrained LLM Services with Authenticity and Reasoning
Jinan Zhou, Rajat Ghosh, Vaishnavi Bhargava +2
When designing LLM services, practitioners care about three key properties: inference-time budget, factual authenticity, and reasoning capacity. However, our analysis shows that no…
Predictive Scaling Laws for Efficient GRPO Training of Large Reasoning Models
Datta Nimmaturi, Vaishnavi Bhargava, Rajat Ghosh +2
Fine-tuning large language models (LLMs) for reasoning tasks using reinforcement learning methods like Group Relative Policy Optimization (GRPO) is computationally expensive. To ad…