8 papers
AgentPatch: Coarse-to-Fine Weak-Task Repair for Merging Agentic Multimodal Large Language Models
Zibo Shao, Baochen Xiong, Chengdong Xu +6
Agentic multimodal large language models (MLLMs) extend multimodal perception and reasoning with planning, tool use, and interaction in dynamic environments. Yet current models are…
Preference-Driven Online Adaptation for Personalized Interaction Initiation in Proactive AI Assistants
Yufeng Wang, Wei Zhang, Zhiquan Wen +5
AI assistants are typically reactive, relying on users to initiate interactions. Proactive assistants go beyond this paradigm by autonomously initiating interactions based on users…
A Survey on the Green Development of Large Models: From Resource-Efficient Architectures to Hardware-Software Co-Design
Linhui Xiao, Guiping Cao, Mingyue Guo +6
The rapid expansion of large-scale AI models has led to significant performance breakthroughs across diverse domains, yet it has also raised critical concerns regarding computation…
RGBT-GroundBench: Visual Grounding Beyond RGB in Complex Real-World Scenarios
Tianyi Zhao, Jiawen Xi, Linhui Xiao +4
Visual grounding (VG) localizes target objects in an image from natural-language expressions. In real-world perception, RGB cues often degrade under low illumination and adverse we…
BARE: Towards Bias-Aware and Reasoning-Enhanced One-Tower Visual Grounding
Hongbing Li, Linhui Xiao, Zihan Zhao +4
Visual Grounding (VG), which aims to locate a specific region referred to by expressions, is a fundamental yet challenging task in the multimodal understanding fields. While recent…
Toward Visual Grounding: A Survey
Linhui Xiao, Xiaoshan Yang, Xiangyuan Lan +2
Visual Grounding, also known as Referring Expression Comprehension and Phrase Grounding, aims to ground the specific region(s) within the image(s) based on the given expression tex…