collaborators

6 papers

cs.CV2026

Closed-Loop Triplet Synergistic Generation for Long-Form Video

Xinlei Yin, Xiulian Peng, Xiao Li +2

Multi-shot long-form video generation remains challenging due to identity drift and compounding inconsistencies across shots. While storyboard-driven pipelines improve controllabil…

cs.CV2026

From Content to Knowledge: Lightning Fast Long-Video Understanding with Neural Knowledge Representations

Yuchen Guan, Xiao Li, Zongyu Guo +4

We propose a new paradigm for long video understanding by treating a long video as a Neural Knowledge Representation (NKR). NKR represents video contents neither as a stream of tok…

cs.CV2026

Hierarchical Long Video Understanding with Audiovisual Entity Cohesion and Agentic Search

Xinlei Yin, Xiulian Peng, Xiao Li +2

Long video understanding presents significant challenges for vision-language models due to extremely long context windows. Existing solutions relying on naive chunking strategies w…

cs.SD2025

Text-Queried Audio Source Separation via Hierarchical Modeling

Xinlei Yin, Xiulian Peng, Xue Jiang +2

Target audio source separation with natural language queries presents a promising paradigm for extracting arbitrary audio events through arbitrary text descriptions. Existing metho…

cs.SD2025

Universal Speech Token Learning via Low-Bitrate Neural Codec and Pretrained Representations

Xue Jiang, Xiulian Peng, Yuan Zhang +1

Current large speech language models are mainly based on semantic tokens from discretization of self-supervised learned representations and acoustic tokens from a neural codec, fol…

cs.CV2025

Bitrate-Controlled Diffusion for Disentangling Motion and Content in Video

Xiao Li, Qi Chen, Xiulian Peng +3

We propose a novel and general framework to disentangle video data into its dynamic motion and static content components. Our proposed method is a self-supervised pipeline with les…