collaborators

13 papers

cs.CV2026

Harnessing Streaming Video in the Wild

Dingyu Yao, Shuhuan Gu, Qingyi Si +8

Vision-Language Models (VLMs) are increasingly required to process unbounded video streams in applications such as video-call assistants, live commentary, and embodied robots. An i…

cs.LG2026

Co-Evolving Policy Distillation

Naibin Gu, Chenxu Yang, Qingyi Si +7

RLVR and OPD have become standard paradigms for post-training. We provide a unified analysis of these two paradigms in consolidating multiple expert capabilities into a single mode…

cs.LG2026

Self-Distilled RLVR

Chenxu Yang, Chuanyu Qin, Qingyi Si +7

On-policy distillation (OPD) has become a popular training paradigm in the LLM community. This paradigm selects a larger model as the teacher to provide dense, fine-grained signals…

cs.SD2026

DegDiT: Controllable Audio Generation with Dynamic Event Graph Guided Diffusion Transformer

Yisu Liu, Chenxing Li, Wanqian Zhang +6

Controllable text-to-audio generation aims to synthesize audio from textual descriptions while satisfying user-specified constraints, including event types, temporal sequences, and…

cs.AI2026

System 1&2 Synergy via Dynamic Model Interpolation

Chenxu Yang, Qingyi Si, Chong Tian +6

Training a unified language model that adapts between intuitive System 1 and deliberative System 2 remains challenging due to interference between their cognitive modes. Recent stu…

cs.AI2025

Test-time Prompt Intervention

Chenxu Yang, Qingyi Si, Mz Dai +5

Test-time compute has led to remarkable success in the large language model (LLM) community, particularly for complex tasks, where longer chains of thought (CoTs) are generated to…