activity
20242026
collaborators

6 papers

cs.MM2026

EditEmoTalk: Controllable Speech-Driven 3D Facial Animation with Continuous Expression Editing

Diqiong Jiang, Kai Zhu, Dan Song +3

Speech-driven 3D facial animation aims to generate realistic and expressive facial motions directly from audio. While recent methods achieve high-quality lip synchronization, they…

cs.LG2025

Anchoring Values in Temporal and Group Dimensions for Flow Matching Model Alignment

Yawen Shao, Jie Xiao, Kai Zhu +4

Group Relative Policy Optimization (GRPO) has proven highly effective in enhancing the alignment capabilities of Large Language Models (LLMs). However, current adaptations of GRPO…

cs.CV2025

State Space Prompting via Gathering and Spreading Spatio-Temporal Information for Video Understanding

Jiahuan Zhou, Kai Zhu, Zhenyu Cui +3

Recently, pre-trained state space models have shown great potential for video classification, which sequentially compresses visual tokens in videos with linear complexity, thereby…

cs.CV2025

FACM: Flow-Anchored Consistency Models

Yansong Peng, Kai Zhu, Yu Liu +4

Continuous-time Consistency Models (CMs) promise efficient few-step generation but face significant challenges with training instability. We argue this instability stems from a fun…

cs.CV2025

Semantically-Aware Game Image Quality Assessment

Kai Zhu, Vignesh Edithal, Le Zhang +2

Assessing the visual quality of video game graphics presents unique challenges due to the absence of reference images and the distinct types of distortions, such as aliasing, textu…

cs.CV2024

GUNet: A Graph Convolutional Network United Diffusion Model for Stable and Diversity Pose Generation

Shuowen Liang, Sisi Li, Qingyun Wang +3

Pose skeleton images are an important reference in pose-controllable image generation. In order to enrich the source of skeleton images, recent works have investigated the generati…