agent evaluation 1asynchronous runtime 1benchmarking 1large language models 1software infrastructure 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.AI2026
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities
Kai Chen, Zichen Ding, Jiaye Ge +22
The paper presents AgentCompass, an open‑source infrastructure that standardizes and simplifies the evaluation of large‑language‑model based autonomous agents by separating benchma…
cs.IR2026
Closing the Auto-Research Loop: An AI Co-Scientist for Production Search Ranking
Liwei Wu, Cho-Jui Hsieh
We present an AI Co-Scientist framework that closes the research loop for the production search-ranking system of a large online travel platform -- pairing LLM agents with direct c…
cs.AI2025
TCIA: A Task-Centric Instruction Augmentation Method for Instruction Finetuning
Simin Ma, Shujian Liu, Jun Tan +7
Diverse instruction data is vital for effective instruction tuning of large language models, as it enables the model to generalize across different types of inputs . Building such…