From the 1 of 9 linked papers with an AI index.
9 papers
NaviDC-OCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
Peng Cai, Zhaofan Zou, Shifa Liu +7
Document parsing aims to transform unstructured documents into structured and machine-readable representations. Recent advances in Vision-Language Models (VLMs) have significantly…
GeoAnchor: Collaborative Reasoning via Latent Decomposition for 3D Spatial Understanding
Hao Li, Han Fang, Zixin Pan +8
GeoAnchor introduces a framework that breaks down 3D spatial information from 2D images into position, direction, and geometry latent components, enabling more accurate and interpr…
Boosting Robust AIGI Detection with LoRA-based Pairwise Training
Ruiyang Xia, Qi Zhang, Yaowen Xu +4
The proliferation of highly realistic AI-Generated Image (AIGI) has necessitated the development of practical detection methods. While current AIGI detectors perform admirably on c…
WAT: Online Video Understanding Needs Watching Before Thinking
Zifan Han, Hongbo Sun, Jinglin Xu +6
Multimodal Large Language Models (MLLMs) have shown strong capabilities in image understanding, motivating recent efforts to extend them to video reasoning. However, existing Video…
RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension
Tianyi Gao, Hao Li, Han Fang +8
Referring Expression Comprehension (REC) is a vision-language task that localizes a specific image region based on a textual description. Existing REC benchmarks primarily evaluate…
Adaptive Evidential Learning for Temporal-Semantic Robustness in Moment Retrieval
Haojian Huang, Kaijing Ma, Jin Chen +8
In the domain of moment retrieval, accurately identifying temporal segments within videos based on natural language queries remains challenging. Traditional methods often employ pr…