collaborators

5 papers

cs.AI2026

LongDocBench: Benchmarking TOC Hierarchy and Contextual Relationship Recovery in Long Documents

Yuefeng Zou, Yichen Lu, Jingxiao Yang +5

Parsing visual documents into machine-readable representations is fundamental to document intelligence. Existing benchmarks focus on page-level element recognition, reading order,…

cs.CV2026

LingDT-VL-OCR: Structure-Aware Document-Level Parsing with Fine-Grained Visual Reference

Siyi Qian, Xiongfei Bai, Bingtao Fu +4

In this paper, we propose LingDT-VL-OCR, a document parsing system tailored to financial-domain documents, transforming ultra-long financial PDFs into semantically consistent, high…

cs.CV2026

AuthFace: Towards Authentic Blind Face Restoration with Face-oriented Generative Diffusion Prior

Guoqiang Liang, Qingnan Fan, Bingtao Fu +3

Blind face restoration (BFR) is a fundamental and challenging problem in computer vision. To faithfully restore high-quality (HQ) photos from poor-quality ones, recent research end…

cs.CV2025

OMGSR: You Only Need One Mid-timestep Guidance for Real-World Image Super-Resolution

Zhiqiang Wu, Zhaomang Sun, Tong Zhou +7

Denoising Diffusion Probabilistic Models (DDPMs) show promising potential in one-step Real-World Image Super-Resolution (Real-ISR). Current one-step Real-ISR methods typically inje…

cs.CV2025

CoMPaSS: Enhancing Spatial Understanding in Text-to-Image Diffusion Models

Gaoyang Zhang, Bingtao Fu, Qingnan Fan +5

Text-to-image (T2I) diffusion models excel at generating photorealistic images but often fail to render accurate spatial relationships. We identify two core issues underlying this…