collaborators

18 papers

cs.CV2026

RoRA: Role-Oriented Regional Allocation for Visual Token Pruning in MLLMs

Qiyanhui Lu, Han Wu, Rongjian Xu +6

Multimodal large language models (MLLMs) encode images as long visual token sequences, making prefilling and KV-cache storage expensive. Existing training-free pruning methods sele…

cs.AI2026

From Question Answering to Task Completion: A Survey on Agent System and Harness Design

Jianyuan Guo, Zhiwei Hao, Chengcheng Wang +14

LLM-based agents mark a shift from passive question answering to active task completion: they perceive environments, invoke tools, maintain state, and act over extended horizons. A…

cs.CL2026

Retrieval-Augmented Linguistic Calibration

Yi-Fan Yeh, Linwei Tao, Minjing Dong +4

Linguistic cues such as "I believe" and "probably" offer an intuitive interface for communicating confidence, yet a generalisable, principled calibration framework for linguistic c…

cs.CV2026

Q-Zoom: Query-Aware Adaptive Perception for Efficient Multimodal Large Language Models

Yuheng Shi, Xiaohuan Pei, Linfeng Wen +2

MLLMs require high-resolution visual inputs for fine-grained tasks like document understanding and dense scene perception. However, current global resolution scaling paradigms indi…

cs.LG2026

Confidence Calibration under Ambiguous Ground Truth

Linwei Tao, Haoyang Luo, Minjing Dong +1

Confidence calibration assumes a unique ground-truth label per input, yet this assumption fails wherever annotators genuinely disagree. Post-hoc calibrators fitted on majority-vote…

cs.CV2026

Mitigating Object Hallucinations in Large Vision-Language Models via Attention Calibration

Younan Zhu, Linwei Tao, Minjing Dong +1

Large Vision-Language Models (LVLMs) exhibit impressive multimodal reasoning capabilities but remain highly susceptible to object hallucination, where models generate responses tha…