activity
20242026
collaborators

7 papers

cs.CV2026

POINTS-Seeker: An Open Recipe for Multimodal Search Agents with Visual Memory Management

Yikun Liu, Yuan Liu, Le Tian +6

Large Multimodal Models (LMMs) excel at visual perception but struggle with real-time, knowledge-intensive queries due to their reliance on static parametric knowledge. While multi…

cs.CV2026

A Sanity Check on Composed Image Retrieval

Yikun Liu, Jiangchao Yao, Weidi Xie +1

Composed Image Retrieval (CIR) aims to retrieve a target image based on a query composed of a reference image, and a relative caption that specifies the desired modification. Despi…

cs.CV2026

POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs

Haicheng Wang, Yuan Liu, Yikun Liu +9

Multimodal Large Language Models (MLLMs) have recently demonstrated remarkable capabilities in cross-modal understanding and generation. However, the rapid growth of visual token s…

cs.CV2026

VersaViT: Enhancing MLLM Vision Backbones via Task-Guided Optimization

Yikun Liu, Yuan Liu, Shangzhe Di +8

Multimodal Large Language Models (MLLMs) have recently achieved remarkable success in visual-language understanding, demonstrating superior high-level semantic alignment within the…

cs.CV2026

POINTS-GUI-G: GUI-Grounding Journey

Zhongyin Zhao, Yuan Liu, Yikun Liu +7

The rapid advancement of vision-language models has catalyzed the emergence of GUI agents, which hold immense potential for automating complex tasks, from online shopping to flight…

cs.CV2025

Learning Streaming Video Representation via Multitask Training

Yibin Yan, Jilan Xu, Shangzhe Di +6

Understanding continuous video streams plays a fundamental role in real-time applications including embodied AI and autonomous driving. Unlike offline video understanding, streamin…