collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2026

FlashSign: Pose-Free Guidance for Efficient Sign Language Video Generation

Liuzhou Zhang, Zeyu Zhang, Biao Wu +10

Sign language plays a crucial role in bridging communication gaps between the deaf and hard-of-hearing communities. However, existing sign language video generation models often re…

cs.CV2026

Twin Co-Adaptive Dialogue for Progressive Image Generation

Jianhui Wang, Yangfan He, Yan Zhong +12

Modern text-to-image generation systems have enabled the creation of remarkably realistic and high-quality visuals, yet they often falter when handling the inherent ambiguities in…

cs.CV2026

PhysicsMind: Sim and Real Mechanics Benchmarking for Physical Reasoning and Prediction in Foundational VLMs and World Models

Chak-Wing Mak, Guanyu Zhu, Boyi Zhang +16

Modern foundational Multimodal Large Language Models (MLLMs) and video world models have advanced significantly in mathematical, common-sense, and visual reasoning, but their grasp…

cs.CV2025

Optimizing Multi-Round Enhanced Training in Diffusion Models for Improved Preference Understanding

Kun Li, Jianhui Wang, Yangfan He +10

Generative AI has significantly changed industries by enabling text-driven image generation, yet challenges remain in achieving high-resolution outputs that align with fine-grained…

cs.CV2025

Efficient Temporal Consistency in Diffusion-Based Video Editing with Adaptor Modules: A Theoretical Framework

Xinyuan Song, Yangfan He, Sida Li +10

Adapter-based methods are commonly used to enhance model performance with minimal additional complexity, especially in video editing tasks that require frame-to-frame consistency.…