3 papers
cs.AI2026
AISE-Bench: A Full-Cycle Curated Benchmark for Information Seeking on Academic Knowledge Graphs
Fanjin Zhang, Zhengyang Wang, Ruixuan Huang +7
Large language models (LLMs) augmented with tools are emerging as autonomous agents capable of using Web engine, APIs, and code to solve complex, long-horizon tasks. Current tool-u…
cs.CL2026
H2HMem: A Multimodal Memory Benchmark for Agents in Human-Human Interactions
Shiping Zhu, Yibo Yang, Zhengyang Wang +3
Large language model agents are increasingly deployed in human-human interaction settings, such as meeting assistants and clinical documentation systems, where they must observe co…
cs.CL2025
CS-Bench: A Comprehensive Benchmark for Large Language Models towards Computer Science Mastery
Xiaoshuai Song, Muxi Diao, Guanting Dong +13
Large language models (LLMs) have demonstrated significant potential in advancing various fields of research and society. However, the current community of LLMs overly focuses on b…