collaborators

19 papers

cs.LG2026

Prefix-Guided On-Policy Distillation: Mining Golden Trajectories from Rollouts

Qingfei Zhao, Huan Song, Shuyu Tian +2

On-policy distillation (OPD) improves reasoning models by applying dense teacher supervision on student-sampled trajectories. However, scaling OPD to long-horizon reasoning exposes…

eess.IV2026

Enhancing Neural Video Compression of Static Scenes with Positive-Incentive Noise

Cheng Yuan, Zhenyu Jia, Jiawei Shao +1

Static scene videos, such as surveillance feeds and videotelephony streams, constitute a dominant share of storage consumption and network traffic. However, both traditional standa…

cs.CL2026

Ruyi2.5 Technical Report

Huan Song, Shuyu Tian, Qingfei Zhao +5

We present Ruyi2.5, a multimodal familial model built on the AI Flow framework. Extending Ruyi2's "Train Once, Deploy Many" paradigm to the multimodal domain, Ruyi2.5 constructs a…

cs.CV2026

Multimodal Continual Learning with MLLMs from Multi-scenario Perspectives

Kai Jiang, Siqi Huang, Xiangyu Chen +4

Multimodal large language models (MLLMs) deployed on devices must adapt to continuously changing visual scenarios such as variations in background and perspective, to effectively p…

cs.LG2026

Controllable Exploration in Hybrid-Policy RLVR for Multi-Modal Reasoning

Zhuoxu Huang, Mengxi Jia, Hao Sun +2

Reinforcement Learning with verifiable rewards (RLVR) has emerged as a primary learning paradigm for enhancing the reasoning capabilities of multi-modal large language models (MLLM…

cs.LG2026

Efficiently Aligning Draft Models via Parameter- and Data-Efficient Adaptation

Luxi Lin, Zhihang Lin, Zhanpeng Zeng +5

Speculative decoding accelerates LLM inference but suffers from performance degradation when target models are fine-tuned for specific domains. A naive solution is to retrain draft…