5 papers
dRAE: Representation Autoencoder with Hyper-Spherical Codes
Tianren Ma, Lin Long, Chuyan Chen +4
In this work, we aim to discretize the high-dimensional visual representations to bridge the gap with language models - a non-trivial challenge, as existing quantization methods su…
AceTone: Bridging Words and Colors for Conditional Image Grading
Tianren Ma, Mingxiang Liao, Xijin Zhang +1
Color affects how we interpret image style and emotion. Previous color grading methods rely on patch-wise recoloring or fixed filter banks, struggling to generalize across creative…
Consolidating Reinforcement Learning for Multimodal Discrete Diffusion Models
Tianren Ma, Mu Zhang, Yibing Wang +1
Optimizing discrete diffusion model (DDM) with rewards remains a challenge: the non-autoregressive paradigm makes importance sampling intractable and rollout complex, puzzling rein…
ReDDiT: Rehashing Noise for Discrete Visual Generation
Tianren Ma, Xiaosong Zhang, Boyu Yang +2
In the visual generative area, discrete diffusion models are gaining traction for their efficiency and compatibility. However, pioneered attempts still fall behind their continuous…
ClawMachine: Learning to Fetch Visual Tokens for Referential Comprehension
Tianren Ma, Lingxi Xie, Yunjie Tian +2
Aligning vision and language concepts at a finer level remains an essential topic of multimodal large language models (MLLMs), particularly for tasks such as referring and groundin…