collaborators

7 papers

cs.CV2026

Towards Flexible, Natural, Efficient Interaction for Conversational Talking Face Generation

Baiqin Wang, Sen Chen, Jiankuo Zhao +3

Conversational talking face generation has recently attracted increasing attention, aiming to synthesize interactive talking videos where characters speak, listen, and respond dyna…

cs.CV2026

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model

Lichen Bai, Tianhao Zhang, Shitong Shao +14

As an increasing majority of global video content is consumed on social platforms for interactive social purposes, video generation models built for social worlds are important but…

cs.CV2026

TurboTalk: Progressive Distillation for One-Step Audio-Driven Talking Avatar Generation

Xiangyu Liu, Feng Gao, Xiaomei Zhang +4

Existing audio-driven video digital human generation models rely on multi-step denoising, resulting in substantial computational overhead that severely limits their deployment in r…

cs.CV2026

Improving Large Vision-Language Models' Understanding for Flow Field Data

Xiaomei Zhang, Hanyu Zheng, Xiangyu Zhu +4

Large Vision-Language Models (LVLMs) have shown impressive capabilities across a range of tasks that integrate visual and textual understanding, such as image captioning and visual…

cs.CE2026

AdaField: Generalizable Surface Pressure Modeling with Physics-Informed Pre-training and Flow-Conditioned Adaptation

Junhong Zou, Wei Qiu, Zhenxu Sun +3

The surface pressure field of transportation systems, including cars, trains, and aircraft, is critical for aerodynamic analysis and design. In recent years, deep neural networks h…

cs.CV2025

Pose-RFT: Enhancing MLLMs for 3D Pose Generation via Hybrid Action Reinforcement Fine-Tuning

Bao Li, Xiaomei Zhang, Miao Xu +3

Generating 3D human poses from multimodal inputs such as images or text requires models to capture both rich spatial and semantic correspondences. While pose-specific multimodal la…