Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation
Lin Shi, Haowei Lin, Zixuan Zhu +123
Evaluating agents on the growing number of agentic benchmarks is challenging because they often require complex environments and agent integrations. We introduce Harbor Adapters, a…
cs.AI2026
STAMP: Provenance-Guided Credit Assignment for Deep Search Agents
Ke Xu, Han Xu, Xinran Chen +6
Reinforcement learning for deep-search agents has largely focused on trajectory-level scoring -- outcome correctness, citation-aware rewards, and evidence coverage. Yet the actions…
cs.AI2026
Data and Evaluation Closed-Loop for Model Capability Enhancement
Zhixuan Li, Jiangan Yuan, Han Xu
Model capability is the central variable in LLM pre-training, yet is never observed directly: data shapes it prospectively, while evaluation reveals it only retrospectively, compre…