From the 2 of 47 linked papers with an AI index.
4 citations · 11 across the 37 of their papers we have counts for
5 papers · 1 filter
Mirror-Fusion Attention for Reflection-Aware Self-Supervised Representation Learning
Ruixin Li, Jin Liu, Yuling Shi +1
Most self-supervised learning (SSL) methods encourage invariance across augmentations, but strict flip invariance can suppress informative left--right correspondences in approximat…
ShredBench: Evaluating the Semantic Reasoning Capabilities of Multimodal LLMs in Document Reconstruction
Zichun Guo, Yuling Shi, Wenhao Zeng +6
Multimodal Large Language Models (MLLMs) have achieved remarkable performance in Visually Rich Document Understanding (VRDU) tasks, but their capabilities are mainly evaluated on p…
Analyzing the Mechanism of Attention Collapse in VGGT from a Dynamics Perspective
Huan Li, Longjun Luo, Yuling Shi +1
Visual Geometry Grounded Transformer (VGGT) delivers state-of-the-art feed-forward 3D reconstruction, yet its global self-attention layer suffers from a drastic collapse phenomenon…
GraphGeo: Multi-Agent Debate Framework for Visual Geo-localization with Heterogeneous Graph Neural Networks
Heng Zheng, Yuling Shi, Xiaodong Gu +6
Visual geo-localization requires extensive geographic knowledge and sophisticated reasoning to determine image locations without GPS metadata. Traditional retrieval methods are con…
Progressive Supernet Training for Efficient Visual Autoregressive Modeling
Xiaoyue Chen, Yuling Shi, Kaiyuan Li +5
Visual Auto-Regressive (VAR) models significantly reduce inference steps through the "next-scale" prediction paradigm. However, progressive multi-scale generation incurs substantia…