2 papers
cs.AI2026
AIDABench: AI Data Analytics Benchmark
Yibo Yang, Fei Lei, Yixuan Sun +24
As AI-driven document understanding and processing tools become increasingly prevalent in real-world applications, the need for rigorous evaluation standards has grown increasingly…
cs.CL2025
Can Large Language Models Play Text Games Well? Current State-of-the-Art and Open Questions
Chen Feng Tsai, Xiaochen Zhou, Sierra S. Liu +3
Large language models (LLMs) such as ChatGPT and GPT-4 have recently demonstrated their remarkable abilities of communicating with human users. In this technical report, we take an…