3 papers
cs.CV2025
UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding
Chenkai Xu, Xu Wang, Zhenyi Liao +3
Consistency models (CMs) have shown promise in the efficient generation of both image and text. This raises the natural question of whether we can learn a unified CM for efficient…
cs.CV2025
Improved Visual-Spatial Reasoning via R1-Zero-Like Training
Zhenyi Liao, Qingsong Xie, Yanhao Zhang +4
Increasing attention has been placed on improving the reasoning capacities of multi-modal large language models (MLLMs). As the cornerstone for AI agents that function in the physi…
cs.CV2024
TLCM: Training-efficient Latent Consistency Model for Image Generation with 2-8 Steps
Qingsong Xie, Zhenyi Liao, Zhijie Deng +2
Distilling latent diffusion models (LDMs) into ones that are fast to sample from is attracting growing research interest. However, the majority of existing methods face two critica…