7 papers
CutVerse: A Compositional GUI Agents Benchmark for Media Post-Production Editing
Haobo Hu, Xiangwu Guo, Zhiheng Chen +4
While GUI agents have made significant progress in web navigation and basic operating system tasks, their capabilities in professional creative workflows remain largely underexplor…
Camera Artist: A Multi-Agent Framework for Cinematic Language Storytelling Video Generation
Haobo Hu, Qi Mao, Yuanhang Li +1
We propose Camera Artist, a multi-agent framework that models a real-world filmmaking workflow to generate narrative videos with explicit cinematic language. While recent multi-age…
Generative Neural Video Compression via Video Diffusion Prior
Qi Mao, Hao Cheng, Tinghan Yang +2
We present GNVC-VD, the first DiT-based generative neural video compression framework built upon an advanced video generation foundation model, where spatio-temporal latent compres…
IC-Effect: Precise and Efficient Video Effects Editing via In-Context Learning
Yuanhang Li, Yiren Song, Junzhe Bai +4
We propose \textbf{IC-Effect}, an instruction-guided, DiT-based framework for few-shot video VFX editing that synthesizes complex effects (\eg flames, particles and cartoon charact…
UniMIC: Token-Based Multimodal Interactive Coding for Human-AI Collaboration
Qi Mao, Tinghan Yang, Jiahao Li +3
The rapid progress of Large Multimodal Models (LMMs) and cloud-based AI agents is transforming human-AI collaboration into bidirectional, multimodal interaction. However, existing…
EmoAgent: A Multi-Agent Framework for Diverse Affective Image Manipulation
Qi Mao, Haobo Hu, Yujie He +3
Affective Image Manipulation (AIM) aims to alter visual elements within an image to evoke specific emotional responses from viewers. However, existing AIM approaches rely on rigid…