Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
TrustDABench: Benchmarking Reliability and Robustness of LLMs for Structured Data Analysis
Boshen Shi, Yize Liu, Chen Zhao +4
LLMs are increasingly used to analyze spreadsheets, CSV files, and other structured data, but producing a correct-looking answer is not the same as producing a trustworthy analysis…
cs.CL2025
TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models
Ce Li, Xiaofan Liu, Zhiyan Song +10
The majority of data in businesses and industries is stored in tables, databases, and data warehouses. Reasoning with table-structured data poses significant challenges for large l…