9 citations · 9 across the 2 of their papers we have counts for
3 papers
cs.LG2026
Group Adaptive Clipping Policy Optimization
Sheng Jia, Xiao Wang, Shiva Prasad Kasiviswanathan +1
Group relative policy optimization for reinforcement learning with verifiable rewards (RLVR) typically uses a fixed importance-sampling (IS) ratio clipping boundary across all roll…
cs.CL2025
Training Large Language Models To Reason In Parallel With Global Forking Tokens
Sheng Jia, Xiao Wang, Shiva Prasad Kasiviswanathan
Although LLMs have demonstrated improved performance by scaling parallel test-time compute, doing so relies on generating reasoning paths that are both diverse and accurate. For ch…
cs.LG2019★ 9 cited
DOM-Q-NET: Grounded RL on Structured Language
Sheng Jia, Jamie Kiros, Jimmy Ba
Building agents to interact with the web would allow for significant improvements in knowledge understanding and representation learning. However, web navigation tasks are difficul…