11 papers
Beyond Static Endpoints: Tool Programs as an Interface for Flexible Agentic Web Services
Mugeng Liu, Shuoqi Li, Yixuan Zhang +1
In the agentic web era, LLM-based agents increasingly invoke web services as tools, yet most interfaces remain \emph{static endpoints} that poorly express long-horizon workflows wi…
SGR-Bench: Benchmarking Search Agents on State-Gated Retrieval
Ningyuan Li, Haiyang Shen, Mugeng Liu +4
Recent advances in large language models and tool-using agents have expanded the range of benchmarked web tasks. Yet an important class of specialized retrieval tasks remains under…
Teaching AI Through Benchmark Construction: QuestBench as a Course-Based Practice for Accountable Knowledge Work
Haiyang Shen, Jiuzheng Wang, Taian Guo +9
As AI becomes part of everyday learning, many courses teach students to use it mainly as a productivity tool: how to prompt, search, summarize, write, code, and use tools more effi…
MindLoom: Composing Thought Modes for Frontier-Level Reasoning Data Synthesis
Haiyang Shen, Taian Guo, Xuanzhong Chen +11
Although LLMs have made substantial progress in reasoning, systematically producing frontier-level reasoning data remains difficult. Existing synthesis methods often have limited v…
DeepWeb-Bench: A Deep Research Benchmark Demanding Massive Cross-Source Evidence and Long-Horizon Derivation
Sixiong Xie, Zhuofan Shi, Haiyang Shen +8
Deep research, in which an agent searches the open web, collects evidence, and derives an answer through extended reasoning, is a prominent use case for frontier language models. F…
From Task to Tutorial: An Automated GUI Framework for Excel Tutorial Document and Video Creation
Yuhang Xie, Jian Mu, Xiaojun Ma +9
Excel is one of the most widely used productivity tools across domains, offering rich functionality but also overwhelming users with its complexity. This creates a persistent deman…