collaborators

6 papers

eess.AS2026

G-STAR: End-to-End Global Speaker-Tracking Attributed Recognition

Jing Peng, Ziyi Chen, Haoyu Li +7

We study timestamped speaker-attributed automatic speech recognition (SA-ASR) for long-form, multi-party speech with overlap. In this setting, chunk-wise inference must preserve me…

cs.CV2026

UniVBench: Towards Unified Evaluation for Video Foundation Models

Jianhui Wei, Xiaotian Zhang, Yichen Li +6

Video foundation models aim to integrate video understanding, generation, editing, and instruction following within a single framework, making them a central direction for next-gen…

cs.CV2025

Input-Aware Sparse Attention for Real-Time Co-Speech Video Generation

Beijia Lu, Ziyi Chen, Jing Xiao +1

Diffusion models can synthesize realistic co-speech video from audio for various applications, such as video creation and virtual agents. However, existing diffusion-based methods…

cs.CV2025

EasyGenNet: An Efficient Framework for Audio-Driven Gesture Video Generation Based on Diffusion Model

Renda Li, Xiaohua Qi, Qiang Ling +4

Audio-driven cospeech video generation typically involves two stages: speech-to-gesture and gesture-to-video. While significant advances have been made in speech-to-gesture generat…

cs.CV2025

Data-free Knowledge Distillation with Diffusion Models

Xiaohua Qi, Renda Li, Long Peng +6

Recently Data-Free Knowledge Distillation (DFKD) has garnered attention and can transfer knowledge from a teacher neural network to a student neural network without requiring any a…

cs.LG2025

SyncDiff: Diffusion-based Talking Head Synthesis with Bottlenecked Temporal Visual Prior for Improved Synchronization

Xulin Fan, Heting Gao, Ziyi Chen +3

Talking head synthesis, also known as speech-to-lip synthesis, reconstructs the facial motions that align with the given audio tracks. The synthesized videos are evaluated on mainl…