2 citations · 2 across the 1 of their papers we have counts for
5 papers
Hierarchical Long Video Understanding with Audiovisual Entity Cohesion and Agentic Search
Xinlei Yin, Xiulian Peng, Xiao Li +2
Long video understanding presents significant challenges for vision-language models due to extremely long context windows. Existing solutions relying on naive chunking strategies w…
Bitrate-Controlled Diffusion for Disentangling Motion and Content in Video
Xiao Li, Qi Chen, Xiulian Peng +3
We propose a novel and general framework to disentangle video data into its dynamic motion and static content components. Our proposed method is a self-supervised pipeline with les…
Text-Queried Audio Source Separation via Hierarchical Modeling
Xinlei Yin, Xiulian Peng, Xue Jiang +2
Target audio source separation with natural language queries presents a promising paradigm for extracting arbitrary audio events through arbitrary text descriptions. Existing metho…
Universal Speech Token Learning via Low-Bitrate Neural Codec and Pretrained Representations
Xue Jiang, Xiulian Peng, Yuan Zhang +1
Current large speech language models are mainly based on semantic tokens from discretization of self-supervised learned representations and acoustic tokens from a neural codec, fol…
Convert and Speak: Zero-shot Accent Conversion with Minimum Supervision
Zhijun Jia, Huaying Xue, Xiulian Peng +1
Low resource of parallel data is the key challenge of accent conversion(AC) problem in which both the pronunciation units and prosody pattern need to be converted. We propose a two…