1 citations · 1 across the 6 of their papers we have counts for
9 papers
Beyond Tables: Doc2DB-Bench for Relationally Faithful Document-to-Database Construction
Zhuowen Liang, Zhengxuan Zhang, Jiayang Wang +2
Practical AI systems increasingly need to turn long, heterogeneous documents into queryable relational databases, not isolated spreadsheets. In domains such as finance, healthcare,…
DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces
Boyan Li, Zhuowen Liang, Yupeng Xie +11
Data agents enable natural-language analytics over organizational workspaces, where relevant evidence may be scattered across databases, structured files, long documents, and multi…
Long-Document QA with Chain-of-Structured-Thought and Fine-Tuned SLMs
Zhuowen Liang, Xiaotian Lin, Zhengxuan Zhang +3
Large language models (LLMs) are widely applied to data analytics over documents, yet direct reasoning over long, noisy documents remains brittle and error-prone. Hence, we study d…
LatentGeo: Learnable Auxiliary Constructions in Latent Space for Multimodal Geometric Reasoning
Haiying Xu, Zihan Wang, Song Dai +3
Despite recent advances in multimodal reasoning, representing auxiliary geometric constructions remains a fundamental challenge for multimodal large language models (MLLMs). Such c…
DocSage: An Information Structuring Agent for Multi-Doc Multi-Entity Question Answering
Teng Lin, Yizhang Zhu, Zhengxuan Zhang +2
Multi-document Multi-entity Question Answering inherently demands models to track implicit logic between multiple entities across scattered documents. However, existing Large Langu…
TableTale: Reviving the Narrative Interplay Between Data Tables and Text in Scientific Papers
Liangwei Wang, Zhengxuan Zhang, Yifan Cao +2
Data tables play a central role in scientific papers. However, their meaning is often co-constructed with surrounding text through narrative interplay, making comprehension cogniti…