2 papers
cs.LG2026
Parameter Exploration for RLVR via Variational Learning
Vatsal Venkatkrishna, Nico Daheim, Iryna Gurevych
Exploration has been a focus of reinforcement learning research for a long time. Recently, there has been growing evidence that it is also an important ingredient in LLM reinforcem…
cs.SE2026
Aletheia: What Makes RLVR For Code Verifiers Tick?
Vatsal Venkatkrishna, Indraneil Paul, Iryna Gurevych
Multi-domain thinking verifiers trained via Reinforcement Learning with Verifiable Rewards (RLVR) are a cornerstone of modern post-training. However, their adoption in code generat…