3 papers
cs.CV2026
dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models
Yi Xin, Siqi Luo, Tianxiang Xu +13
Diffusion Multi-modal Large Language Models (dMLLMs) have recently emerged as a novel architecture unifying image generation and understanding. However, developing effective and ef…
cs.CV2025
PMMD: A pose-guided multi-view multi-modal diffusion for person generation
Ziyu Shang, Haoran Liu, Rongchao Zhang +2
Generating consistent human images with controllable pose and appearance is essential for applications in virtual try on, image editing, and digital human creation. Current methods…
cs.CV2025
Low-Cost Test-Time Adaptation for Robust Video Editing
Jianhui Wang, Yinda Chen, Yangfan He +6
Video editing is a critical component of content creation that transforms raw footage into coherent works aligned with specific visual and narrative objectives. Existing approaches…