3 papers
cs.CV2026
Uni-ViGU: Towards Unified Video Generation and Understanding via A Diffusion-Based Video Generator
Luozheng Qin, Jia Gong, Qian Qiao +6
Unified multimodal models integrating visual understanding and generation face a fundamental challenge: visual generation incurs substantially higher computational costs than under…
cs.CV2026
TextFlux: An OCR-Free DiT Model for High-Fidelity Multilingual Scene Text Synthesis
Yu Xie, Jielei Zhang, Pengyu Chen +5
Diffusion-based scene text synthesis has progressed rapidly, yet existing methods commonly rely on additional visual conditioning modules and require large-scale annotated data to…
cs.CV2025
Temporal-Conditional Referring Video Object Segmentation with Noise-Free Text-to-Video Diffusion Model
Ruixin Zhang, Jiaqing Fan, Yifan Liao +2
Referring Video Object Segmentation (RVOS) aims to segment specific objects in a video according to textual descriptions. We observe that recent RVOS approaches often place excessi…