collaborators

36 papers

cs.CV2026

When Vision Becomes Text: Visual Token Pruning via Cross-Modal Residual Guidance in VLMs

Congyang Ou, Ruike Song, Yang Zhou +3

Abundant visual information strengthens vision-language model (VLM) perception, yet massive visual tokens raise inference costs. Existing visual token pruning methods rely on simil…

cs.CV2026

DeltaV: Thinking with Visual State Updates in Unified Large Multimodal Models

Pengjie Wang, Linger Deng, Zujia Zhang +6

Current Unified Large Multimodal Models (ULMMs) support interleaved multimodal reasoning through textual reasoning and intermediate visual states, but typically generate each visua…

cs.AI2026

EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language Models

Yiyang Fang, Wenke Huang, Pei Fu +5

Multimodal Large Language Models (MLLMs) have shown remarkable progress in visual reasoning and understanding tasks but still struggle to capture the complexity and subjectivity of…

cs.IR2026

ELVA: Exploring Ranking-Driven Universal Multimodal Retrieval

Yuhan Liu, Pei Fu, Hang Li +8

Leveraging Multimodal Large Language Models (MLLMs) via contrastive learning has become a mainstream paradigm for improving the performance of Universal Multimodal Retrieval (UMR).…

cs.AI2026

Xiaomi-GUI-0 Technical Report

Wanxia Cao, Chengzhen Duan, Pei Fu +29

Graphical user interface (GUI) agents build on vision-language models to complete user tasks end-to-end in real applications through interface actions such as tapping, swiping, tex…

cs.AI2026

GAIA: A Data Flywheel System for Training GUI Test-Time Scaling Critic Models

Shaokang Wang, Pei Fu, Ruoceng Zhang +7

While Large Vision-Language Models (LVLMs) have significantly advanced GUI agents' capabilities in parsing textual instructions, interpreting screen content, and executing tasks, a…