6 papers
SimRecon: SimReady Compositional Scene Reconstruction from Real Videos
Chong Xia, Kai Zhu, Zizhuo Wang +3
Compositional scene reconstruction seeks to create object-centric representations rather than holistic scenes from real-world videos, which is natively applicable for simulation an…
EditEmoTalk: Controllable Speech-Driven 3D Facial Animation with Continuous Expression Editing
Diqiong Jiang, Kai Zhu, Dan Song +3
Speech-driven 3D facial animation aims to generate realistic and expressive facial motions directly from audio. While recent methods achieve high-quality lip synchronization, they…
FACM: Flow-Anchored Consistency Models
Yansong Peng, Kai Zhu, Yu Liu +4
Continuous-time Consistency Models (CMs) promise efficient few-step generation but face significant challenges with training instability. We argue this instability stems from a fun…
Anchoring Values in Temporal and Group Dimensions for Flow Matching Model Alignment
Yawen Shao, Jie Xiao, Kai Zhu +4
Group Relative Policy Optimization (GRPO) has proven highly effective in enhancing the alignment capabilities of Large Language Models (LLMs). However, current adaptations of GRPO…
State Space Prompting via Gathering and Spreading Spatio-Temporal Information for Video Understanding
Jiahuan Zhou, Kai Zhu, Zhenyu Cui +3
Recently, pre-trained state space models have shown great potential for video classification, which sequentially compresses visual tokens in videos with linear complexity, thereby…
Semantically-Aware Game Image Quality Assessment
Kai Zhu, Vignesh Edithal, Le Zhang +2
Assessing the visual quality of video game graphics presents unique challenges due to the absence of reference images and the distinct types of distortions, such as aliasing, textu…