activity
20232026
collaborators

5 papers

cs.CV2026

LLaVAFlow: Preserving Latent Alignment Flow for Parameter-Efficient Multimodal Fine-Tuning

Muyao Yuan, Muyan Jiao, Jiangyong Ying +5

While Multimodal Large Language Models (MLLMs) exhibit strong generalization, visual instruction tuning for downstream tasks inevitably causes catastrophic forgetting, impairing ov…

cs.CV2026

Mitigating Entangled Steering in Large Vision-Language Models for Hallucination Reduction

Yuanhong Zhang, Zhaoyang Wang, Xin Zhang +2

Large Vision-Language Models (LVLMs) have achieved remarkable success across cross-modal tasks but remain hindered by hallucinations, producing textual outputs inconsistent with vi…

cs.CV2025

InfoCLIP: Bridging Vision-Language Pretraining and Open-Vocabulary Semantic Segmentation via Information-Theoretic Alignment Transfer

Muyao Yuan, Yuanhong Zhang, Weizhan Zhang +4

Recently, the strong generalization ability of CLIP has facilitated open-vocabulary semantic segmentation, which labels pixels using arbitrary text. However, existing methods that…

cs.CV2025

InfoSAM: Fine-Tuning the Segment Anything Model from An Information-Theoretic Perspective

Yuanhong Zhang, Muyao Yuan, Weizhan Zhang +4

The Segment Anything Model (SAM), a vision foundation model, exhibits impressive zero-shot capabilities in general tasks but struggles in specialized domains. Parameter-efficient f…

cs.AI2023

Data Quality-aware Mixed-precision Quantization via Hybrid Reinforcement Learning

Yingchun Wang, Jingcai Guo, Song Guo +1

Mixed-precision quantization mostly predetermines the model bit-width settings before actual training due to the non-differential bit-width sampling process, obtaining sub-optimal…