collaborators

6 papers

cs.CV2026

A Visual Dependence-Aware Framework for Multimodal Unsupervised Continual Post-Training

Kaichen Li, Zhilin Zhu, Jianhao Huang +7

In this paper, we explore a novel task of Multimodal Unsupervised Continual Post-Training (MU-CPT), enabling deployed MLLMs to continually evolve from streaming unlabeled data. Exi…

cs.AI2026

AgentPatch: Coarse-to-Fine Weak-Task Repair for Merging Agentic Multimodal Large Language Models

Zibo Shao, Baochen Xiong, Chengdong Xu +6

Agentic multimodal large language models (MLLMs) extend multimodal perception and reasoning with planning, tool use, and interaction in dynamic environments. Yet current models are…

cs.HC2026

Preference-Driven Online Adaptation for Personalized Interaction Initiation in Proactive AI Assistants

Yufeng Wang, Wei Zhang, Zhiquan Wen +5

AI assistants are typically reactive, relying on users to initiate interactions. Proactive assistants go beyond this paradigm by autonomously initiating interactions based on users…

cs.LG2026

A Survey on the Green Development of Large Models: From Resource-Efficient Architectures to Hardware-Software Co-Design

Linhui Xiao, Guiping Cao, Mingyue Guo +6

The rapid expansion of large-scale AI models has led to significant performance breakthroughs across diverse domains, yet it has also raised critical concerns regarding computation…

cs.CV2026

BARE: Towards Bias-Aware and Reasoning-Enhanced One-Tower Visual Grounding

Hongbing Li, Linhui Xiao, Zihan Zhao +4

Visual Grounding (VG), which aims to locate a specific region referred to by expressions, is a fundamental yet challenging task in the multimodal understanding fields. While recent…

cs.CV2024

OneRef: Unified One-tower Expression Grounding and Segmentation with Mask Referring Modeling

Linhui Xiao, Xiaoshan Yang, Fang Peng +2

Constrained by the separate encoding of vision and language, existing grounding and referring segmentation works heavily rely on bulky Transformer-based fusion en-/decoders and a v…