activity
20242026
collaborators

14 papers

cs.LG2026

QUADS: Stabilizing NVFP4 Reinforcement Learning for MoE via QUantization-error Alignment across Dual Sides

Zhengyang Zhuge, Hao Yu, Xin Wang +4

Rollout generation is a major bottleneck in Reinforcement Learning (RL) for Mixture-of-Experts (MoE) Large Language Models, motivating low-precision rollout acceleration such as FP…

cs.LG2026

Survive or Collapse: The Asymmetric Roles of Data Gating and Reward Grounding in Self-Play RL

Sophia Xiao Pu, Zhaotian Weng, Chengzhi Liu +4

Self-play reinforcement learning trains language models on their own generated tasks, co-evolving a proposer and solver without human labels. Recent systems report strong reasoning…

cs.AI2026

ECPO: Evidence-Coupled Policy Optimization for Evidence-Certified Candidate Ranking

Miaobo Hu, Shuhao Hu, BoKun Wang +5

Ranking systems used in decision-support settings should not only order candidates but also expose evidence that can be independently checked. We study evidence-certified candidate…

cs.CL2026

Towards Context-Invariant Safety Alignment for Large Language Models

Yixu Wang, Yang Yao, Xin Wang +4

Preference-based post-training aligns LLMs with human intent, yet safety behavior often remains brittle. A model may refuse a harmful request in a standard prompt but comply when t…

cs.LG2026

AGPO: Adaptive Group Policy Optimization with Dual Statistical Feedback

Miaobo Hu, Shuhao Hu, Bokun Wang +5

Reinforcement learning improves LLM reasoning, but PPO/GRPO typically use fixed clipping and decoding temperature, which makes training brittle and tuning-heavy. We propose Adaptiv…

cs.CV2026

SAVER: Selective As-Needed Vision Evidence for Multimodal Information Extraction

Miaobo Hu, Shuhao Hu, Bokun Wang +5

Multimodal IE in social media is difficult because a post may attach multiple images that are weakly related, redundant, or even misleading with respect to the text. In this settin…