activity
20212026
most citedKimi K2.5: Visual Agentic Intelligence

2 citations · 3 across the 15 of their papers we have counts for

collaborators
Showing cs.CVShow all

18 papers · 1 filter

cs.CV2026

State-Conditioned Visual Evidence Retrieval for Fine-Grained Perception in Document Vision-Language Models

Mingxu Chai, Chenyu Liu, Ziyu Shen +7

Compared with typical vision-language tasks, document parsing places stronger demands on fine-grained visual perception. Existing vision-language model (VLM)-based parsing approach…

cs.CV2026

Hierarchical Evidence-Driven Reasoning for Long Document Understanding

Junyu Xiong, Yonghui Wang, Rongjian Gu +4

Retrieval-Augmented Generation (RAG) streamlines long-document understanding by leveraging retrieval mechanisms to restrict input images to a highly curated subset. However, existi…

cs.CV2026

CAGE-SGG: Counterfactual Active Graph Evidence for Open-Vocabulary Scene Graph Generation

Suiyang Guang, Chenyu Liu, Ruohan Zhang +1

Open-vocabulary scene graph generation (SGG) aims to describe visual scenes with flexible and fine-grained relation phrases beyond a fixed predicate vocabulary. While recent vision…

cs.CV2026

Prefix-Adaptive Block Diffusion for Efficient Document Recognition

Mingxu Chai, Ziyu Shen, Chenyu Liu +3

Block Diffusion Models (BDMs) support parallel generation, flexible-length output, and KV caching, making them promising for efficient document parsing. However, existing BDMs bind…

cs.CV2026

TDATR: Improving End-to-End Table Recognition via Table Detail-Aware Learning and Cell-Level Visual Alignment

Chunxia Qin, Chenyu Liu, Pengcheng Xia +4

Tables are pervasive in diverse documents, making table recognition (TR) a fundamental task in document analysis. Existing modular TR pipelines separately model table structure and…

cs.CV2025

CaliTex: Geometry-Calibrated Attention for View-Coherent 3D Texture Generation

Chenyu Liu, Hongze Chen, Jingzhi Bao +7

Despite major advances brought by diffusion-based models, current 3D texture generation systems remain hindered by cross-view inconsistency -- textures that appear convincing from…