2 papers
cs.CV2026
Dreaming in Flow: Generative Grounding Feedback for Self-Evolving Unified Multimodal Models
Ke Hao, Yuanzhi Liang, Tingxi Chen +5
Unified multimodal models integrate visual understanding and generation within a single network, yet the two capabilities are commonly optimized as separate tasks. We introduce Gen…
cs.SD2026
Enhancing Flow Matching with A Unified Guidance Framework for Efficient and Robust Speech Synthesis
Zuda Yu, Qianhui Xu, Ting Chen +5
Flow Matching (FM) has emerged as a powerful paradigm for speech generation but remains constrained by high inference latency and timbre leakage. To address these bottlenecks, we p…