collaborators

6 papers

cs.LG2026

I-SDPO: Instance-Level Adaptive Self-Distillation Policy Optimization

Yubo Zhang, Xinhong Ma, Zezhong Tan +1

Group Relative Policy Optimization (GRPO) learns from reward differences within a rollout group, but receives no useful relative signal when every sampled response is incorrect. Pr…

cs.LG2026

Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective

Feng Zhang, Xinhong Ma, Ziqiang Dong +5

Group Relative Policy Optimization (GRPO) is one of the most widely adopted RLVR algorithms for post-training large language models on reasoning tasks. We first show that GRPO admi…

cs.CV2026

RAVE: Re-Allocating Visual Attention in Large Multimodal Models

Xi Leng, Xinhong Ma, Ziqiang Dong +4

Large multimodal models (LMMs) inherit the self-attention mechanism of pretrained language backbones, yet standard attention can exhibit suboptimal allocation, including cross-moda…

cs.LG2026

More Edits, More Stable: Understanding the Lifelong Normalization in Sequential Model Editing

Xin Ma, Wei Chen, Qi Liu +4

Lifelong Model Editing aims to continuously update evolving facts in Large Language Models while preserving unrelated knowledge and general capabilities, yet it remains plagued by…

cs.CV2026

ADHint: Adaptive Hints with Difficulty Priors for Reinforcement Learning

Feng Zhang, Zezhong Tan, Xinhong Ma +6

To address the limited capability expansion and low sample efficiency of Reinforcement Learning (RL), recent methods have integrated ''hints'' into post-training, which are prefix…

cs.AI2025

Towards Flash Thinking via Decoupled Advantage Policy Optimization

Zezhong Tan, Hang Gao, Xinhong Ma +2

Recent Large Reasoning Models (LRMs) have achieved remarkable performance in solving complex problems via supervised fine-tuning (SFT) and reinforcement learning (RL). Although exi…