most citedMotion Control for Enhanced Complex Action Video Generation

1 citations · 1 across the 4 of their papers we have counts for

collaborators

5 papers

cs.CV2025

ReWatch-R1: Boosting Complex Video Reasoning in Large Vision-Language Models through Agentic Data Synthesis

Congzhi Zhang, Zhibin Wang, Yinchao Ma +5

While Reinforcement Learning with Verifiable Reward (RLVR) significantly advances image reasoning in Large Vision-Language Models (LVLMs), its application to complex video reasonin…

cs.CV2025

Raccoon: Multi-stage Diffusion Training with Coarse-to-Fine Curating Videos

Zhiyu Tan, Junyan Wang, Hao Yang +4

Text-to-video generation has demonstrated promising progress with the advent of diffusion models, yet existing approaches are limited by dataset quality and computational resources…

cs.CV20241 cited

Motion Control for Enhanced Complex Action Video Generation

Qiang Zhou, Shaofeng Zhang, Nianzu Yang +2

Existing text-to-video (T2V) models often struggle with generating videos with sufficiently pronounced or complex actions. A key limitation lies in the text prompt's inability to p…

cs.CV2024

Multimodal LLM Enhanced Cross-lingual Cross-modal Retrieval

Yabing Wang, Le Wang, Qiang Zhou +4

Cross-lingual cross-modal retrieval (CCR) aims to retrieve visually relevant content based on non-English queries, without relying on human-labeled cross-modal data pairs during tr…

cs.CV2024

I2EBench: A Comprehensive Benchmark for Instruction-based Image Editing

Yiwei Ma, Jiayi Ji, Ke Ye +6

Significant progress has been made in the field of Instruction-based Image Editing (IIE). However, evaluating these models poses a significant challenge. A crucial requirement in t…