3 papers
cs.CV2026
dRAE: Representation Autoencoder with Hyper-Spherical Codes
Tianren Ma, Lin Long, Chuyan Chen +4
In this work, we aim to discretize the high-dimensional visual representations to bridge the gap with language models - a non-trivial challenge, as existing quantization methods su…
cs.AI2025
Consolidating Reinforcement Learning for Multimodal Discrete Diffusion Models
Tianren Ma, Mu Zhang, Yibing Wang +1
Optimizing discrete diffusion model (DDM) with rewards remains a challenge: the non-autoregressive paradigm makes importance sampling intractable and rollout complex, puzzling rein…
cs.CV2025
CC-Diff: Enhancing Contextual Coherence in Remote Sensing Image Synthesis
Mu Zhang, Yunfan Liu, Yue Liu +2
Existing image synthesis methods for natural scenes focus primarily on foreground control, often reducing the background to simplistic textures. Consequently, these approaches tend…