30 papers
Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation
Chi Kit Wong, Ye Pan, Yuanhuiyi Lyu +6
The paper proposes Ego Scene Augmentation (ESA), a framework that uses an Ego-element Graph to improve the spatial perception of multimodal large language models for egocentric vis…
Perceptual Flow Network for Visually Grounded Reasoning
Yangfu Li, Yuning Gong, Hongjian Zhan +8
Despite the success of Large-Vision Language Models (LVLMs), general optimization objectives (e.g., standard MLE) fail to constrain visual trajectories, leading to language bias an…
Are We Using the Right Benchmark: An Evaluation Framework for Visual Token Compression Methods
Chenfei Liao, Wensong Wang, Zichen Wen +10
Recent efforts to accelerate inference in Multimodal Large Language Models (MLLMs) have largely focused on visual token compression. The effectiveness of these methods is commonly…
TC-AE: Unlocking Token Capacity for Deep Compression Autoencoders
Teng Li, Ziyuan Huang, Cong Chen +5
We propose TC-AE, a ViT-based architecture for deep compression autoencoders. Existing methods commonly increase the channel number of latent representations to maintain reconstruc…
Unlocking Multimodal Document Intelligence: From Current Triumphs to Future Frontiers of Visual Document Retrieval
Yibo Yan, Jiahao Huo, Guanbo Feng +12
With the rapid proliferation of multimodal information, Visual Document Retrieval (VDR) has emerged as a critical frontier in bridging the gap between unstructured visually rich da…
SAP: Segment Any 4K Panorama
Lutao Jiang, Zidong Cao, Weikai Chen +14
Promptable instance segmentation is widely adopted in embodied and AR systems, yet the performance of foundation models trained on perspective imagery often degrades on 360° panor…