collaborators

5 papers

cs.MM2026

AcoustiTrace: When Plausible Sound Violates Physics

Shiyang Li, Yuewen Cao, Yihao Liu +4

Recent audio-video generators can produce semantically plausible and apparently synchronized sound, yet may still violate the acoustic processes implied by visible events and envir…

cs.RO2026

GenTrack: Physical Alignment for Robot-Native Motion Generation and Zero-Shot Humanoid Tracking

Zeyu Ling, Xinyao Yu, Renye Yan +4

General-purpose humanoid trackers can execute diverse references, but their zero-shot coverage depends on large embodied corpora that are costly to extend. Text-to-motion generator…

cs.CV2026

PRISM: Streaming Human Motion Generation with Per-Joint Latent Decomposition

Zeyu Ling, Qing Shuai, Teng Zhang +3

Text-to-motion generation has advanced with larger corpora and stronger generators, yet many models still rely on holistic frame- or clip-level latents that entangle trajectory, or…

cs.AI2026

SyncLipMAE: Contrastive Masked Pretraining for Audio-Visual Talking-Face Representation

Zeyu Ling, Xiaodong Gu, Jiangnan Tang +1

We introduce SyncLipMAE, a self-supervised pretraining framework for talking-face video that learns synchronization-aware and transferable facial dynamics from unlabeled audio-visu…

cs.CV2025

VersatileMotion: A Unified Framework for Motion Synthesis and Comprehension

Zeyu Ling, Bo Han, Shiyang Li +3

Large language models (LLMs) are, by design, inherently capable of multi-task learning: through a unified next-token prediction paradigm, they can naturally address a wide variety…