4 papers
TRACE-Bench: Decomposing and Diagnosing Multi-Reference Image Generation
Haoran Wang, Chaofan Ma, Ran Yi +1
Despite recent advances in unified multimodal models for multi-reference image generation, existing benchmarks remain organized around predefined task types (e.g., "subject composi…
VicEdit: Learning to Edit Videos from Visual In-Context Examples
Yuji Wang, Teng Hu, Yuheng Chen +6
Despite progress in instruction-based video editing, unimodal textual instructions inherently struggle to convey fine-grained textures and complex dynamics. To bridge this perceptu…
PersonaShot: Benchmarking Person-Centric Narrative Continuity in Multi-Shot Video Generation
Yuji Wang, Yuheng Chen, Teng Hu +7
Video generation is rapidly evolving from single-shot clips to multi-shot narratives, where the human character serves as the core narrative anchor. However, existing benchmarks ma…
Spatial Temporal Synergy: Balancing Change and Invariance in Text Driven 3D Human Motion Editing
Shaohui Lin, Zhenwu Shi, Jingyu Gong +5
Text-driven human motion editing aims to modify existing motion sequences according to natural language instructions while maintaining the structural consistency of the original mo…