8 papers
VicEdit: Learning to Edit Videos from Visual In-Context Examples
Yuji Wang, Teng Hu, Yuheng Chen +6
Despite progress in instruction-based video editing, unimodal textual instructions inherently struggle to convey fine-grained textures and complex dynamics. To bridge this perceptu…
PersonaShot: Benchmarking Person-Centric Narrative Continuity in Multi-Shot Video Generation
Yuji Wang, Yuheng Chen, Teng Hu +7
Video generation is rapidly evolving from single-shot clips to multi-shot narratives, where the human character serves as the core narrative anchor. However, existing benchmarks ma…
In-Context Forcing: Uncovering Context Effects in Autoregressive Video Diffusion
Lingxiao Yang, Liu Liu, Moran Li +4
Current few-step autoregressive video diffusion models depend on previous fully denoised clean frames as context for all denoising steps of the current frame. However, these clean…
CRAFT: Clinical Reward-Aligned Finetuning for Medical Image Synthesis
Yunsung Chung, Alex El Darzi, Carlo El Khoury +3
Foundation diffusion models can generate photorealistic natural images, but adapting them to medical imaging remains challenging. In medical adaptation, limited labeled data can ex…
CoDCL: Counterfactual-Inspired Augmentation Contrastive Learning for Temporal Link Prediction in Social Networks
Hantong Feng, Duxin Chen, Wenwu Yu
Temporal link prediction is crucial for rapidly growing social networks. Existing methods often overlook the underlying causal mechanisms that drive link formation, making it diffi…
Soul: Breathe Life into Digital Human for High-fidelity Long-term Multimodal Animation
Jiangning Zhang, Junwei Zhu, Zhenye Gan +14
We propose a multimodal-driven framework for high-fidelity long-term digital human animation termed , which generates semantically coherent videos from a single-fram…