works on

From the 1 of 16 linked papers with an AI index.

most citedAutoResearchBench: Benchmarking AI Agents on Complex Scientific Literature Discovery

1 citations · 1 across the 8 of their papers we have counts for

collaborators
Showing cs.AIShow all

8 papers · 1 filter

cs.AI2026

From Relevance to Execution Utility: Reward-Aware Dynamic Execution Gating for Skill-Based LLM Agents

Liang He, Jingbo Wen, Hongyu Gu +5

Agent skills are increasingly used to equip large language model (LLM) agents with reusable procedural knowledge. Although recent work has substantially improved skill retrieval du…

cs.AI2026

AgentPanel: Toward a New Paradigm for Human--AI Collaboration in Exploring Scientific Questions

Zhiyao Cui, Qianyi Wang, Haoyang Yan +26

Identifying promising scientific ideas remains an important challenge in research practice. Researchers commonly rely on small-group discussions or one-to-one interactions with a s…

cs.AI2026

SKT: Skill-Use Training at Scale via Verified Synthetic Data Generation

Zelin Tan, Yiqun Zhang, Hao Li +11

Agent skills have become an important mechanism for equipping language-model agents with reusable procedural knowledge. However, providing skills alone does not guarantee that curr…

cs.AI20261 cited

AutoResearchBench: Benchmarking AI Agents on Complex Scientific Literature Discovery

Lei Xiong, Kun Luo, Ziyi Xia +15

Autonomous scientific research is significantly advanced thanks to the development of AI agents. One key step in this process is finding the right scientific literature, whether to…

cs.AI2026

Beyond Gemini-3-Pro: Revisiting LLM Routing and Aggregation at Scale

Shengji Tang, Weihao Lin, Peng Ye +9

Large Language Models (LLMs) have rapidly advanced, with Gemini-3-Pro setting a new performance milestone. In this work, we explore collective intelligence as an alternative to mon…

cs.AI2026

LLMRouterBench: A Massive Benchmark and Unified Framework for LLM Routing

Hao Li, Yiqun Zhang, Zhaoyan Guo +9

Large language model (LLM) routing assigns each query to the most suitable model from an ensemble. We introduce LLMRouterBench, a large-scale benchmark and unified framework for LL…