1 citations · 1 across the 3 of their papers we have counts for
4 papers
Stellar: Scalable Multimodal Document Retrieval for Natural Language Queries
Yuxiang Guo, Zhonghao Hu, Yuren Mao +5
Multimodal document retrieval--selecting the most relevant multimodal document from a large corpus to answer a natural language query--plays an essential role in Retrieval-Augmente…
DTBench: A Synthetic Benchmark for Document-to-Table Extraction
Yuxiang Guo, Zhuoran Du, Nan Tang +3
Document-to-table (Doc2Table) extraction derives structured tables from unstructured documents under a target schema, enabling reliable and verifiable SQL-based data analytics. Alt…
PiLLar: Matching for Pivot Table Schema via LLM-guided Monte-Carlo Tree Search
Yunjun Gao, Chuangyu Ouyang, Congcong Ge +1
Pivot tables are ubiquitous in data lakes of modern data ecosystems, making accurate schema matching over pivot tables a key prerequisite for data integration. In this paper, we fo…
MontePrep: Monte-Carlo-Driven Automatic Data Preparation without Target Data Instances
Congcong Ge, Yachuan Liu, Yixuan Tang +3
In commercial systems, a pervasive requirement for automatic data preparation (ADP) is to transfer relational data from disparate sources to targets with standardized schema specif…