2 papers
cs.LG2026
Selective-Advantage Entropy-Adaptive Horizon GRPO: Asymmetric Token-Level Discounting for Efficient Reinforcement Learning of Language Models
Chirag Chawla, Rohan Charudatt Salvi, Madhav S. Baidya
Group Relative Policy Optimisation (GRPO) has emerged as an effective reinforcement-learning algorithm for aligning language models on reasoning tasks, but it treats every token po…
cs.CL2025
PERCS: Persona-Guided Controllable Biomedical Summarization Dataset
Rohan Charudatt Salvi, Chirag Chawla, Dhruv Jain +3
Automatic medical text simplification plays a key role in improving health literacy by making complex biomedical research accessible to diverse readers. However, most existing reso…