4 papers
SyncLipMAE: Contrastive Masked Pretraining for Audio-Visual Talking-Face Representation
Zeyu Ling, Xiaodong Gu, Jiangnan Tang +1
We introduce SyncLipMAE, a self-supervised pretraining framework for talking-face video that learns synchronization-aware and transferable facial dynamics from unlabeled audio-visu…
EPIC: Efficient Prompt Interaction for Text-Image Classification
Xinyao Yu, Hao Sun, Zeyu Ling +5
In recent years, large-scale pre-trained multimodal models (LMMs) generally emerge to integrate the vision and language modalities, achieving considerable success in multimodal tas…
NeutronTP: Load-Balanced Distributed Full-Graph GNN Training with Tensor Parallelism
Xin Ai, Hao Yuan, Zeyu Ling +6
Graph neural networks (GNNs) have emerged as a promising direction. Training large-scale graphs that relies on distributed computing power poses new challenges. Existing distribute…
VersatileMotion: A Unified Framework for Motion Synthesis and Comprehension
Zeyu Ling, Bo Han, Shiyang Li +3
Large language models (LLMs) are, by design, inherently capable of multi-task learning: through a unified next-token prediction paradigm, they can naturally address a wide variety…