1 citations · 1 across the 5 of their papers we have counts for
11 papers
InstanceAnimator: Multi-Instance Sketch Video Colorization
Yinhan Zhang, Yue Ma, Bingyuan Wang +5
We propose InstanceAnimator, a novel Diffusion Transformer framework for multi-instance sketch video colorization. Existing methods suffer from three core limitations: inflexible u…
Manifold-Aware Exploration for Reinforcement Learning in Video Generation
Mingzhe Zheng, Weijie Kong, Yue Wu +9
Group Relative Policy Optimization (GRPO) methods for video generation like FlowGRPO remain far less reliable than their counterparts for language models and images. This gap arise…
LongVideoAgent: Multi-Agent Reasoning with Long Videos
Runtao Liu, Ziyi Liu, Jiaqi Tang +4
Recent advances in multimodal LLMs and systems that use tools for long-video QA point to the promise of reasoning over hour-long episodes. However, many methods still compress cont…
Follow-Your-Shape: Shape-Aware Image Editing via Trajectory-Guided Region Control
Zeqian Long, Mingzhe Zheng, Kunyu Feng +6
While recent flow-based image editing models demonstrate general-purpose capabilities across diverse tasks, they often struggle to specialize in challenging scenarios -- particular…
MagicAnime: A Hierarchically Annotated, Multimodal and Multitasking Dataset with Benchmarks for Cartoon Animation Generation
Shuolin Xu, Bingyuan Wang, Zeyu Cai +5
Generating high-quality cartoon animations multimodal control is challenging due to the complexity of non-human characters, stylistically diverse motions and fine-grained emotions.…
Calligrapher: Freestyle Text Image Customization
Yue Ma, Qingyan Bai, Hao Ouyang +8
We introduce Calligrapher, a novel diffusion-based framework that innovatively integrates advanced text customization with artistic typography for digital calligraphy and design ap…