6 papers
HiEviDR-Bench: A Benchmark for Hierarchical Evidence Aggregation in Deep Research
Yubo Sun, Chunyi Peng, Yukun Yan +6
Deep research requires models to retrieve, connect, and synthesize evidence from large-scale heterogeneous sources to answer complex queries and produce analytical reports. Existin…
MinerU-Popo: Universal Post-Processing Model for Structured Document Parsing
Bangrui Xu, Ziyang Miao, Xuanhe Zhou +7
VLM-based OCR models have become the de facto choice for document parsing, as they can accurately extract page-level elements (e.g., paragraphs within individual pages) together wi…
MoDora: Tree-Based Semi-Structured Document Analysis System
Bangrui Xu, Qihang Yao, Zirui Tang +8
Semi-structured documents integrate diverse interleaved data elements (e.g., tables, charts, hierarchical paragraphs) arranged in various and often irregular layouts. These documen…
MinerU2.5-Pro: Pushing the Limits of Data-Centric Document Parsing at Scale
Bin Wang, Tianyao He, Linke Ouyang +40
Current document parsing methods advance primarily through model architecture innovation, while systematic engineering of training data remains underexplored. Yet state-of-the-art…
Qute: Towards Quantum-Native Database
Muzhi Chen, Xuanhe Zhou, Wei Zhou +7
This paper envisions a quantum database (Qute) that treats quantum computation as a first-class execution option. Unlike prior simulation-based methods that either run quantum algo…
LLM/Agent-as-Data-Analyst: A Survey
Zirui Tang, Weizheng Wang, Zihang Zhou +16
Large language models (LLMs) and agent techniques have brought a fundamental shift in the functionality and development paradigm of data analysis tasks (a.k.a LLM/Agent-as-Data-Ana…