5 papers
DocTrace: Towards Traceable Long Document VQA via Hierarchical Evidence Graph Reasoning
Le Xiang, Zhicheng Guan, Hong Chen +5
Long Document Visual Question Answering (LongDocVQA) requires Multimodal Large Language Models (MLLMs) to locate, integrate, and reason over heterogeneous document elements distrib…
HPD-Parsing: Hierarchical Parallel Document Parsing
Shu Wei, Jingjing Wu, Lingshu Zhang +10
Efficient teamwork typically combines global coordination with parallel execution, a principle not yet fully reflected in unified Vision-Language Model (VLM)-based document parsers…
P-MTP: Efficient Document Parsing via Multi-Token Prediction with Progressive Depth Scaling
Le Xiang, Chenxi Zhai, Shu Wei +5
Vision-Language Models (VLMs) have revolutionized document parsing by enabling end-to-end mapping from images to structured text, imposing a significant latency bottleneck, particu…
HEad and neCK TumOR (HECKTOR) 2025: Benchmark of Segmentation, Diagnosis, and Prognosis in Multimodal PET/CT
Numan Saeed, Salma Hassan, Shahad Hardan +27
Head and neck cancers (HNC) represent a significant global health burden, with accurate tumor delineation being essential for effective radiotherapy planning. The complexity of the…
autoPET IV challenge: Incorporating organ supervision and human guidance for lesion segmentation in PET/CT
Junwei Huang, Yingqi Hao, Yitong Luo +6
Lesion Segmentation in PET/CT scans is an essential part of modern oncological workflows. To address the challenges of time-intensive manual annotation and high inter-observer vari…