collaborators

5 papers

cs.AI2026

DocTrace: Towards Traceable Long Document VQA via Hierarchical Evidence Graph Reasoning

Le Xiang, Zhicheng Guan, Hong Chen +5

Long Document Visual Question Answering (LongDocVQA) requires Multimodal Large Language Models (MLLMs) to locate, integrate, and reason over heterogeneous document elements distrib…

cs.CL2026

HPD-Parsing: Hierarchical Parallel Document Parsing

Shu Wei, Jingjing Wu, Lingshu Zhang +10

Efficient teamwork typically combines global coordination with parallel execution, a principle not yet fully reflected in unified Vision-Language Model (VLM)-based document parsers…

cs.CV2026

P-MTP: Efficient Document Parsing via Multi-Token Prediction with Progressive Depth Scaling

Le Xiang, Chenxi Zhai, Shu Wei +5

Vision-Language Models (VLMs) have revolutionized document parsing by enabling end-to-end mapping from images to structured text, imposing a significant latency bottleneck, particu…

cs.CV2026

HEad and neCK TumOR (HECKTOR) 2025: Benchmark of Segmentation, Diagnosis, and Prognosis in Multimodal PET/CT

Numan Saeed, Salma Hassan, Shahad Hardan +27

Head and neck cancers (HNC) represent a significant global health burden, with accurate tumor delineation being essential for effective radiotherapy planning. The complexity of the…

eess.IV2025

autoPET IV challenge: Incorporating organ supervision and human guidance for lesion segmentation in PET/CT

Junwei Huang, Yingqi Hao, Yitong Luo +6

Lesion Segmentation in PET/CT scans is an essential part of modern oncological workflows. To address the challenges of time-intensive manual annotation and high inter-observer vari…