activity
20242026
collaborators

20 papers

cs.CV2026

NaviDC-OCR: Navigating Document Parsing Across Digital and Camera-Captured Documents

Peng Cai, Zhaofan Zou, Shifa Liu +7

Document parsing aims to transform unstructured documents into structured and machine-readable representations. Recent advances in Vision-Language Models (VLMs) have significantly…

cs.CV2026

Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval

Zelong Sun, Jun Wang, Kaicheng Yang +3

Unified multimodal retrieval aims to identify candidates that satisfy complex user intent expressed through heterogeneous inputs. Although Large Vision-Language Model (LVLM)-based…

cs.CV2026

Efficient, Validation-Free Intrinsic Quality Estimation for Large-Scale Face Recognition Datasets

Zhichao Chen, Yongle Zhao, Kaicheng Yang +3

We propose Intrinsic Quality (IQ), a validation-free metric designed to estimate the inherent potential of face recognition (FR) datasets to produce high-performance models without…

cs.CV2026

HoloQ-VLA: Uniform W4A4 Quantization of Vision-Language-Action Models

Xinyu Wang, Mingze Li, Sicheng Lyu +6

Vision-Language-Action (VLA) models unify perception, reasoning, and control in a single policy, but their multi-billion-parameter backbones and diffusion-based action heads make o…

cs.CV2026

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence

Xiang An, Yin Xie, Feilong Tang +27

We introduce LLaVA-OneVision-2 (LLaVA-OV-2), the most capable vision-language model in the LLaVA-OneVision series to date, achieving superior performance across a broad range of mu…

cs.IR2026

FD-RAG: Federated Dual-System Retrieval-Augmented Generation

Tianhao Gao, Kai Yang, Yiyang Li

Retrieval-augmented generation (RAG) has emerged as a paradigm for grounding large language models in external knowledge, yet most existing RAG systems assume centralized knowledge…