27 papers
DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models
ZhiYan Hou, Xinyu Tang, Hongyan An +9
Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models using automatically verifiable outcome signals, but these signals…
Continual Learning in Transition
Zhiyan Hou, Dan Zhang, Tao Feng +11
Classical continual learning (CL) has primarily focused on enabling models to update and retain knowledge through parameter-centric mechanisms, e.g., training strategies, architect…
PixVL: Self-Supervised Training of Pixel-Level MLLMs via a Unified Mask--Text Consistency Cycle
Yicheng Xiao, Haoxuan Ma, Caorui Li +7
Recent studies develop pixel-level multimodal large language models (MLLMs) that support both Region Segmentation and Region Understanding, extending multimodal interaction from wh…
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL
Ruiming Liang, Yi Zhong, Yizhen Yuan +6
Modern large language models (LLMs) are expected not just to answer correctly, but to adapt their behavior to different human values and use cases. As a result, multi-reward reinfo…
ReLoop-UME: Recurrent Depth with Learnable Retrieval Registers for Universal Multimodal Embedding
Shijie Wang, Xiangzhao Hao, Yueti Li +3
Universal multimodal embedding (UME) maps heterogeneous multimodal inputs into a shared embedding space. Existing UME models either form embeddings through single forward encoding…
On-Policy Distillation with Curriculum Turn-level Guidance for Multi-turn Agents
Gengsheng Li, Mao Zheng, Mingyang Song +8
Multi-turn agents that plan, invoke tools, and interact with environments offer a promising paradigm for solving complex tasks, yet their capabilities typically rely on very large…