12 citations · 12 across the 9 of their papers we have counts for
5 papers · 1 filter
TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
Peng Cai, Zhaofan Zou, Shifa Liu +7
Document parsing aims to transform unstructured documents into structured and machine-readable representations. Recent advances in Vision-Language Models (VLMs) have significantly…
LAGS: Low-Altitude Gaussian Splatting with Groupwise Heterogeneous Graph Learning
Yikun Wang, Yujie Wan, Wei Zuo +4
Low-altitude Gaussian splatting (LAGS) facilitates 3D scene reconstruction by aggregating aerial images from distributed drones. However, as LAGS prioritizes maximizing reconstruct…
GeoVista: Web-Augmented Agentic Visual Reasoning for Geolocalization
Yikun Wang, Zuyan Liu, Ziyi Wang +3
Current research on agentic visual reasoning enables deep multimodal understanding but primarily focuses on image manipulation tools, leaving a gap toward more general-purpose agen…
MoIIE: Mixture of Intra- and Inter-Modality Experts for Large Vision Language Models
Dianyi Wang, Siyuan Wang, Zejun Li +6
Large Vision-Language Models (LVLMs) have demonstrated remarkable performance across multi-modal tasks by scaling model size and training data. However, these dense LVLMs incur sig…
Autoregressive Semantic Visual Reconstruction Helps VLMs Understand Better
Dianyi Wang, Wei Song, Yikun Wang +4
Typical large vision-language models (LVLMs) apply autoregressive supervision solely to textual sequences, without fully incorporating the visual modality into the learning process…