Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Beyond Rephrasing: Book-Level Organization Improves Synthetic Textbook Data for Mid-Training
Jiawen Tao, Miao Peng, Yaoming Li +7
Synthetic textbook data has improved language model pre-training, but prior work largely treats the benefit as a property of generated content or local rewriting style. We study a…
cs.AI2026
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios
Weihuang Zheng, Tianyuan Zou, Eileen Ye +5
Large Language Models (LLMs) are increasingly deployed as agents that interact with stateful environments over multiple steps: gathering hidden information, composing tool calls, a…
cs.AI2026
Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic AI
Jiaqi Shao, Hanck Chen, Wei Zhang +2
Agent benchmarks increasingly evaluate repository editing, web research, terminal use, and long-horizon interaction. Their scores support capability claims only when the evaluation…