collaborators

7 papers

cs.CV2026

Latent Noise Mask for Reducing Visual Redundancy in Multimodal Large Language Models

Kai Jiang, Ruishu Zhu, Siqi Huang +2

Multimodal large language models (MLLMs) often fail in fine-grained visual reasoning, as question-relevant visual cues are diluted by dense and redundant image tokens. Recent multi…

cs.CV2026

Multimodal Continual Learning with MLLMs from Multi-scenario Perspectives

Kai Jiang, Siqi Huang, Xiangyu Chen +4

Multimodal large language models (MLLMs) deployed on devices must adapt to continuously changing visual scenarios such as variations in background and perspective, to effectively p…

cs.CV2025

Laytrol: Preserving Pretrained Knowledge in Layout Control for Multimodal Diffusion Transformers

Sida Huang, Siqi Huang, Ping Luo +1

With the development of diffusion models, enhancing spatial controllability in text-to-image generation has become a vital challenge. As a representative task for addressing this c…

cs.LG2025

CoLM: Collaborative Large Models via A Client-Server Paradigm

Siqi Huang, Sida Huang, Hongyuan Zhang

Large models have achieved remarkable performance across a range of reasoning and understanding tasks. Prior work often utilizes model ensembles or multi-agent systems to collabora…

cs.AI2025

AI Flow: Perspectives, Scenarios, and Approaches

Hongjun An, Wenhan Hu, Sida Huang +11

Pioneered by the foundational information theory by Claude Shannon and the visionary framework of machine intelligence by Alan Turing, the convergent evolution of information and c…

cs.LG2025

Learn Beneficial Noise as Graph Augmentation

Siqi Huang, Yanchen Xu, Hongyuan Zhang +1

Although graph contrastive learning (GCL) has been widely investigated, it is still a challenge to generate effective and stable graph augmentations. Existing methods often apply h…