6 papers
G-STAR: End-to-End Global Speaker-Tracking Attributed Recognition
Jing Peng, Ziyi Chen, Haoyu Li +7
We study timestamped speaker-attributed automatic speech recognition (SA-ASR) for long-form, multi-party speech with overlap. In this setting, chunk-wise inference must preserve me…
UniVBench: Towards Unified Evaluation for Video Foundation Models
Jianhui Wei, Xiaotian Zhang, Yichen Li +6
Video foundation models aim to integrate video understanding, generation, editing, and instruction following within a single framework, making them a central direction for next-gen…
Input-Aware Sparse Attention for Real-Time Co-Speech Video Generation
Beijia Lu, Ziyi Chen, Jing Xiao +1
Diffusion models can synthesize realistic co-speech video from audio for various applications, such as video creation and virtual agents. However, existing diffusion-based methods…
EasyGenNet: An Efficient Framework for Audio-Driven Gesture Video Generation Based on Diffusion Model
Renda Li, Xiaohua Qi, Qiang Ling +4
Audio-driven cospeech video generation typically involves two stages: speech-to-gesture and gesture-to-video. While significant advances have been made in speech-to-gesture generat…
Data-free Knowledge Distillation with Diffusion Models
Xiaohua Qi, Renda Li, Long Peng +6
Recently Data-Free Knowledge Distillation (DFKD) has garnered attention and can transfer knowledge from a teacher neural network to a student neural network without requiring any a…
SyncDiff: Diffusion-based Talking Head Synthesis with Bottlenecked Temporal Visual Prior for Improved Synchronization
Xulin Fan, Heting Gao, Ziyi Chen +3
Talking head synthesis, also known as speech-to-lip synthesis, reconstructs the facial motions that align with the given audio tracks. The synthesized videos are evaluated on mainl…