5 papers
Beyond Tables: Doc2DB-Bench for Relationally Faithful Document-to-Database Construction
Zhuowen Liang, Zhengxuan Zhang, Jiayang Wang +2
Practical AI systems increasingly need to turn long, heterogeneous documents into queryable relational databases, not isolated spreadsheets. In domains such as finance, healthcare,…
DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces
Boyan Li, Zhuowen Liang, Yupeng Xie +11
Data agents enable natural-language analytics over organizational workspaces, where relevant evidence may be scattered across databases, structured files, long documents, and multi…
QwenPaw-Data: Bridging Facts, Methodology, and Execution for Autonomous Enterprise Data Analytics
Tianjing Zeng, Yuntao Hong, Zhongjun Ding +17
Enterprise data analysis is emerging as a distinct frontier for autonomous agents. Compared with general-purpose interaction and software engineering, it operates in an open, ambig…
Long-Document QA with Chain-of-Structured-Thought and Fine-Tuned SLMs
Zhuowen Liang, Xiaotian Lin, Zhengxuan Zhang +3
Large language models (LLMs) are widely applied to data analytics over documents, yet direct reasoning over long, noisy documents remains brittle and error-prone. Hence, we study d…
DataPuzzle: Breaking Free from the Hallucinated Promise of LLMs in Data Analysis
Zhengxuan Zhang, Zhuowen Liang, Yin Wu +3
Large language models (LLMs) are increasingly applied to multi-modal data analysis -- not necessarily because they offer the most precise answers, but because they provide fluent,…