works on

From the 2 of 30 linked papers with an AI index.

activity
20242026
collaborators

30 papers

cs.CV2026

ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding

Mingkang Dong, Muxin Pu, Jie Li +8

ObjectStream introduces a training‑free method that extracts latent objects from frozen Video‑LLM representations and uses them as persistent memory anchors to improve streaming vi…

cs.CV2026

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation

Chi Kit Wong, Ye Pan, Yuanhuiyi Lyu +6

The paper proposes Ego Scene Augmentation (ESA), a framework that uses an Ego-element Graph to improve the spatial perception of multimodal large language models for egocentric vis…

cs.CV2026

OmniCoT: A Benchmark for Global and Multi-Step Panoramic Reasoning

Haocong He, Chenfei Liao, Zichen Wen +13

Multimodal Large Language Models (MLLMs) have demonstrated promising spatial reasoning capabilities, while these abilities remain underexplored in the emerging visual modality of p…

cs.CV2026

Seg-Agent: Test-Time Multimodal Reasoning for Training-Free Language-Guided Segmentation

Chao Hao, Jun Xu, Ji Du +6

Language-guided segmentation transcends the scope limitations of traditional semantic segmentation, enabling models to segment arbitrary target regions based on natural language in…

cs.CV2026

Are We Using the Right Benchmark: An Evaluation Framework for Visual Token Compression Methods

Chenfei Liao, Wensong Wang, Zichen Wen +10

Recent efforts to accelerate inference in Multimodal Large Language Models (MLLMs) have largely focused on visual token compression. The effectiveness of these methods is commonly…

cs.CL2026

Unlocking Multimodal Document Intelligence: From Current Triumphs to Future Frontiers of Visual Document Retrieval

Yibo Yan, Jiahao Huo, Guanbo Feng +12

With the rapid proliferation of multimodal information, Visual Document Retrieval (VDR) has emerged as a critical frontier in bridging the gap between unstructured visually rich da…