2 papers
cs.CV2026
Forge-and-Quench: Enhancing Image Generation for Higher Fidelity in Unified Multimodal Models
Yanbing Zeng, Jia Wang, Hanghang Ma +4
Integrating image generation and understanding into a single framework has become a pivotal goal in the multimodal domain. However, how understanding can effectively assist generat…
cs.LG2025
Generalization in VAE and Diffusion Models: A Unified Information-Theoretic Analysis
Qi Chen, Jierui Zhu, Florian Shkurti
Despite the empirical success of Diffusion Models (DMs) and Variational Autoencoders (VAEs), their generalization performance remains theoretically underexplored, especially lackin…