4 papers
TRACE-Bench: Decomposing and Diagnosing Multi-Reference Image Generation
Haoran Wang, Chaofan Ma, Ran Yi +1
Despite recent advances in unified multimodal models for multi-reference image generation, existing benchmarks remain organized around predefined task types (e.g., "subject composi…
VicEdit: Learning to Edit Videos from Visual In-Context Examples
Yuji Wang, Teng Hu, Yuheng Chen +6
Despite progress in instruction-based video editing, unimodal textual instructions inherently struggle to convey fine-grained textures and complex dynamics. To bridge this perceptu…
PersonaShot: Benchmarking Person-Centric Narrative Continuity in Multi-Shot Video Generation
Yuji Wang, Yuheng Chen, Teng Hu +7
Video generation is rapidly evolving from single-shot clips to multi-shot narratives, where the human character serves as the core narrative anchor. However, existing benchmarks ma…
Efficient Audio-Visual Generation via Synchrony-Aware Cross-Modal Sparse Attention
Shengchuan Gao, Teng Hu, Bohao Feng +4
Recent audio-visual generation models can synthesize synchronized video and sound in a unified diffusion process, but their inference cost remains high because long video token seq…