3 papers
cs.AI2026
ScholarGym: Benchmarking Large Language Model Capabilities in the Information-Gathering Stage of Deep Research
Hao Shen, Hang Yang, Zhouhong Gu +1
Large language models have advanced from single-turn question answering to deep research systems that iteratively decompose research questions, invoke retrieval tools, and synthesi…
cs.SI2025
ComGPT: Detecting Local Community Structure with Large Language Models
Li Ni, Haowen Shen, Lin Mu +2
Large Language Models (LLMs), like GPT-3.5-turbo, have demonstrated the ability to understand graph structures and have achieved excellent performance in various graph reasoning ta…
cs.CL2024
CIBench: Evaluating Your LLMs with a Code Interpreter Plugin
Chuyu Zhang, Songyang Zhang, Yingfan Hu +8
While LLM-Based agents, which use external tools to solve complex problems, have made significant progress, benchmarking their ability is challenging, thereby hindering a clear und…