collaborators

11 papers

cs.AI2026

DAREBench: Deployment-Aware and Reliable Evaluation of Models as Agents

Yu Liu, Zhilin Liu, Zhiwei Yang +7

As large language models evolve from question-answering systems into general-purpose agents, evaluation must move beyond static answer correctness to assess multimodal perception,…

cs.AI2026

Making Every Tool Call Count: Necessary Tool-Evidence Path Rewards for Agentic Vision-Language Models

Xingming Long, Yu Liu, Zhiwei Yang +7

Modern vision-language models (VLMs) can directly answer many image-grounded questions, yet they often struggle with complex queries requiring fine-grained visual details or extern…

cs.CV2026

Switch-Reasoner: Learn When to Think in Multitask Mixtures via Reinforcement Learning

Yiyang Fang, Pei Fu, Jinjie Li +7

Multimodal Large Language Models (MLLMs) often follow a fixed Think-then-Answer paradigm, which is inefficient in heterogeneous multitask settings because simple inputs may not req…

cs.CV2026

DeltaV: Thinking with Visual State Updates in Unified Large Multimodal Models

Pengjie Wang, Linger Deng, Zujia Zhang +6

Current Unified Large Multimodal Models (ULMMs) support interleaved multimodal reasoning through textual reasoning and intermediate visual states, but typically generate each visua…

cs.AI2026

Xiaomi-GUI-0 Technical Report

Wanxia Cao, Chengzhen Duan, Pei Fu +29

Graphical user interface (GUI) agents build on vision-language models to complete user tasks end-to-end in real applications through interface actions such as tapping, swiping, tex…

cs.CV2026

UniTranslator: A Unified Multi-modal Framework for End-to-end In-Image Machine Translation

Jiahao Lyu, Pei Fu, Zhenhang Li +6

In-Image Machine Translation (IIMT) aims to translate scene text in an image and render the translated text back into the original regions while preserving the overall visual appea…