6 papers
Latent Noise Mask for Reducing Visual Redundancy in Multimodal Large Language Models
Kai Jiang, Ruishu Zhu, Siqi Huang +2
Multimodal large language models (MLLMs) often fail in fine-grained visual reasoning, as question-relevant visual cues are diluted by dense and redundant image tokens. Recent multi…
Batch Loss Score for Dynamic Data Pruning
Qing Zhou, Bingxuan Zhao, Tao Yang +3
Dynamic data pruning accelerates deep learning by selectively omitting less informative samples during training. While per-sample loss is a common importance metric, obtaining it c…
Explore How to Inject Beneficial Noise in MLLMs
Ruishu Zhu, Sida Huang, Ziheng Jiao +1
Multimodal Large Language Models (MLLMs) have played an increasingly important role in multimodal intelligence. However, the existing fine-tuning methods often ignore cross-modal h…
Rectified Noise: A Generative Model Using Positive-incentive Noise
Zhenyu Gu, Yanchen Xu, Sida Huang +2
Rectified Flow (RF) has been widely used as an effective generative model. Although RF is primarily based on probability flow Ordinary Differential Equations (ODE), recent studies…
Laytrol: Preserving Pretrained Knowledge in Layout Control for Multimodal Diffusion Transformers
Sida Huang, Siqi Huang, Ping Luo +1
With the development of diffusion models, enhancing spatial controllability in text-to-image generation has become a vital challenge. As a representative task for addressing this c…
CoLM: Collaborative Large Models via A Client-Server Paradigm
Siqi Huang, Sida Huang, Hongyuan Zhang
Large models have achieved remarkable performance across a range of reasoning and understanding tasks. Prior work often utilizes model ensembles or multi-agent systems to collabora…