activity
20202026
most citedAssessing Hidden Risks of LLMs: An Empirical Study on Robustness, Consistency, and Credibility

14 citations · 82 across the 47 of their papers we have counts for

collaborators
Showing 2026 · cs.LGShow all

5 papers · 2 filters

cs.LG2026

DE-Venus: A Data-Efficient RLVR Framework for Large Language Models

Shenzhi Yang, Guangcheng Zhu, Kai Tang +11

Reinforcement learning with verifiable rewards (RLVR) improves large language model reasoning, but its practical scaling is constrained by expensive on-policy rollouts and the cost…

cs.LG2026

Frozen Cores Need Task Signal: Fisher-Whitened Cross-Covariance for Low-Resource LLM Adaptation

Wentao Ye, Zhanming Shen, Zhiqing Xiao +3

Parameter-efficient fine-tuning is usually framed as a question of how many parameters to update. Under a severe trainable-state budget, however, where those coefficients act is eq…

cs.LG2026

CriPO: Enhancing Rubric-based RL via Self-Distillation

Mingxuan Xia, Yuhang Yang, Chao Ye +7

Rubric-based RL has recently shown promise in improving LLMs on open-ended tasks. A widely recognized limitation of rubric-based RL is limited exploration: criteria that no rollout…

cs.LG2026

OPRD: On-Policy Representation Distillation

Shenzhi Yang, Guangcheng Zhu, Bowen Song +8

On-policy distillation (OPD) supervises the student exclusively in the output space by matching next-token distributions. This paradigm suffers from two limitations: (i) a high-var…

cs.LG2026

FastBUS: A Fast Bayesian Framework for Unified Weakly-Supervised Learning

Ziquan Wang, Haobo Wang, Ke Chen +2

Machine Learning often involves various imprecise labels, leading to diverse weakly supervised settings. While recent methods aim for universal handling, they usually suffer from c…