8 citations · 14 across the 19 of their papers we have counts for
10 papers · 1 filter
DE-Venus: A Data-Efficient RLVR Framework for Large Language Models
Shenzhi Yang, Guangcheng Zhu, Kai Tang +11
Reinforcement learning with verifiable rewards (RLVR) improves large language model reasoning, but its practical scaling is constrained by expensive on-policy rollouts and the cost…
Frozen Cores Need Task Signal: Fisher-Whitened Cross-Covariance for Low-Resource LLM Adaptation
Wentao Ye, Zhanming Shen, Zhiqing Xiao +3
Parameter-efficient fine-tuning is usually framed as a question of how many parameters to update. Under a severe trainable-state budget, however, where those coefficients act is eq…
GeoMin: Data-Efficient Semi-Supervised RLVR via Geometric Distribution Modeling
Guangcheng Zhu, Shenzhi Yang, Haobo Wang +9
Reinforcement learning with verifiable rewards (RLVR) significantly advances LLM reasoning, yet it faces a dilemma: standard supervised scaling is throttled by high annotation cost…
Smart Picks in the Dark: Towards Efficient RLVR for Reasoning via Tracing Metacognitive Pivots
Guangcheng Zhu, Shenzhi Yang, Haobo Wang +7
Reinforcement learning with verifiable rewards (RLVR) has greatly advanced large reasoning models (LRMs), but it requires timely training on a huge fully-annotated dataset. To this…
Can LLMs Learn to Reason Robustly under Noisy Supervision?
Shenzhi Yang, Guangcheng Zhu, Bowen Song +7
Reinforcement Learning with Verifiable Rewards (RLVR) effectively trains reasoning models that rely on abundant perfect labels, but its vulnerability to unavoidable noisy labels du…
FastBUS: A Fast Bayesian Framework for Unified Weakly-Supervised Learning
Ziquan Wang, Haobo Wang, Ke Chen +2
Machine Learning often involves various imprecise labels, leading to diverse weakly supervised settings. While recent methods aim for universal handling, they usually suffer from c…