collaborators

6 papers

cs.CV2026

Latent Noise Mask for Reducing Visual Redundancy in Multimodal Large Language Models

Kai Jiang, Ruishu Zhu, Siqi Huang +2

Multimodal large language models (MLLMs) often fail in fine-grained visual reasoning, as question-relevant visual cues are diluted by dense and redundant image tokens. Recent multi…

cs.LG2026

Batch Loss Score for Dynamic Data Pruning

Qing Zhou, Bingxuan Zhao, Tao Yang +3

Dynamic data pruning accelerates deep learning by selectively omitting less informative samples during training. While per-sample loss is a common importance metric, obtaining it c…

cs.CV2025

Explore How to Inject Beneficial Noise in MLLMs

Ruishu Zhu, Sida Huang, Ziheng Jiao +1

Multimodal Large Language Models (MLLMs) have played an increasingly important role in multimodal intelligence. However, the existing fine-tuning methods often ignore cross-modal h…

cs.LG2025

Rectified Noise: A Generative Model Using Positive-incentive Noise

Zhenyu Gu, Yanchen Xu, Sida Huang +2

Rectified Flow (RF) has been widely used as an effective generative model. Although RF is primarily based on probability flow Ordinary Differential Equations (ODE), recent studies…

cs.CV2025

Laytrol: Preserving Pretrained Knowledge in Layout Control for Multimodal Diffusion Transformers

Sida Huang, Siqi Huang, Ping Luo +1

With the development of diffusion models, enhancing spatial controllability in text-to-image generation has become a vital challenge. As a representative task for addressing this c…

cs.LG2025

CoLM: Collaborative Large Models via A Client-Server Paradigm

Siqi Huang, Sida Huang, Hongyuan Zhang

Large models have achieved remarkable performance across a range of reasoning and understanding tasks. Prior work often utilizes model ensembles or multi-agent systems to collabora…