4 papers
Representation Learning in Diffusion and Flow-based Model: An Application Aspect
Yanchen Xu, Sida Huang, Zhenyu Gu +3
Diffusion models and flow-based models have recently become the dominant paradigms in generative modeling, largely due to their ability to learn rich, multi-level visual representa…
Latent Noise Mask for Reducing Visual Redundancy in Multimodal Large Language Models
Kai Jiang, Ruishu Zhu, Siqi Huang +2
Multimodal large language models (MLLMs) often fail in fine-grained visual reasoning, as question-relevant visual cues are diluted by dense and redundant image tokens. Recent multi…
ViewMask-1-to-3: Multi-View Consistent Image Generation via Multimodal Discrete Diffusion Models
Ruishu Zhu, Zhihao Huang, Jiacheng Sun +3
Motivated by discrete diffusion's success in language-vision modeling, we explore its potential for multi-view generation, a task dominated by continuous approaches. We introduce V…
Explore How to Inject Beneficial Noise in MLLMs
Ruishu Zhu, Sida Huang, Ziheng Jiao +1
Multimodal Large Language Models (MLLMs) have played an increasingly important role in multimodal intelligence. However, the existing fine-tuning methods often ignore cross-modal h…