works on

From the 1 of 11 linked papers with an AI index.

collaborators
Showing cs.AIShow all

5 papers · 1 filter

cs.AI2026

Beyond Rephrasing: Book-Level Organization Improves Synthetic Textbook Data for Mid-Training

Jiawen Tao, Miao Peng, Yaoming Li +7

The paper introduces a pipeline that creates synthetic textbooks by clustering source material, planning hierarchical tables of contents, and assembling sections into full books, s…

cs.AI2026

Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents

Bowen Ye, Rang Li, Qibin Yang +10

Large language models are increasingly deployed as autonomous agents for multi-step workflows in real-world software environments. However, existing agent benchmarks are limited by…

cs.AI2026

MCPAgentBench: A Real-world Task Benchmark for Evaluating LLM Agent MCP Tool Use

Wenrui Liu, Zixiang Liu, Elsie Dai +5

Large Language Models (LLMs) are increasingly serving as autonomous agents, and their utilization of external tools via the Model Context Protocol (MCP) is considered a future tren…

cs.AI2026

ARC: Active and Reflection-driven Context Management for Long-Horizon Information Seeking Agents

Yilun Yao, Shan Huang, Elsie Dai +5

Large language models are increasingly deployed as research agents for deep search and long-horizon information seeking, yet their performance often degrades as interaction histori…

cs.AI2026

EntroCoT: Enhancing Chain-of-Thought via Adaptive Entropy-Guided Segmentation

Zihang Li, Yuhang Wang, Yikun Zong +6

Chain-of-Thought (CoT) prompting has significantly enhanced the mathematical reasoning capabilities of Large Language Models. We find existing fine-tuning datasets frequently suffe…