3 papers
cs.AI2026
AIDABench: AI Data Analytics Benchmark
Yibo Yang, Fei Lei, Yixuan Sun +24
As AI-driven document understanding and processing tools become increasingly prevalent in real-world applications, the need for rigorous evaluation standards has grown increasingly…
cs.AI2024
DeepDiveAI: Identifying AI Related Documents in Large Scale Literature Data
Zhou Xiaochen, Liang Xingzhou, Zou Hui +2
In this paper, we propose a method to automatically classify AI-related documents from large-scale literature databases, leading to the creation of an AI-related literature dataset…
cs.CL2024
MANGO: A Benchmark for Evaluating Mapping and Navigation Abilities of Large Language Models
Peng Ding, Jiading Fang, Peng Li +6
Large language models such as ChatGPT and GPT-4 have recently achieved astonishing performance on a variety of natural language processing tasks. In this paper, we propose MANGO, a…