collaborators

7 papers

cs.AI2026

TOPS: First-Principles Visual Token Pruning via Constructing Token Optimal Preservation Sets for Efficient MLLM Inference

Tinghao Wang, Yichen Guo, Rui Huang +11

Multimodal large language models (MLLMs) have achieved strong multimodal reasoning capabilities, but their efficiency is limited by the large number of visual tokens, which introdu…

cs.SD2026

HeadRouter: Dynamic Head-Weight Routing for Task-Adaptive Audio Token Pruning in Large Audio Language Models

Peize He, Yaodi Luo, Xiaoqian Liu +7

Recent large audio language models (LALMs) demonstrate remarkable capabilities in processing extended multi-modal sequences, yet incur high inference costs. Token compression is an…

cs.CV2026

Models as Lego Builders: Assembling Malice from Benign Blocks via Semantic Blueprints

Chenxi Li, Xianggan Liu, Dake Shen +9

Despite the rapid progress of Large Vision-Language Models (LVLMs), the integration of visual modalities introduces new safety vulnerabilities that adversaries can exploit to elici…

cs.CV2026

MAP: Mitigating Hallucinations in Large Vision-Language Models with Map-Level Attention Processing

Chenxi Li, Yichen Guo, Benfang Qian +5

Large Vision-Language Models (LVLMs) have achieved impressive performance in multimodal tasks, but they still suffer from hallucinations, i.e., generating content that is grammatic…

cs.CV2025

MagicWand: A Universal Agent for Generation and Evaluation Aligned with User Preference

Zitong Xu, Dake Shen, Yaosong Du +3

Recent advances in AIGC (Artificial Intelligence Generated Content) models have enabled significant progress in image and video generation. However, users still struggle to obtain…

cs.LG2025

Medical priority fusion: achieving dual optimization of sensitivity and interpretability in nipt anomaly detection

Xiuqi Ge, Zhibo Yao, Yaosong Du

Clinical machine learning faces a critical dilemma in high-stakes medical applications: algorithms achieving optimal diagnostic performance typically sacrifice the interpretability…