collaborators

16 papers

cs.AI2026

Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence

Brian Wang, Bin Feng, Xiaoman Pan +26

Apollo did not reach the Moon merely because its engineers could solve difficult equations. It succeeded by turning a distant ambition into a mission architecture of explicit objec…

cs.CL2026

LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers

Tao Feng, Fangxu Yu, Haozhen Zhang +9

No single large language model (LLM) is optimal across all queries and budget constraints, making model routing essential for cost-effective deployment. Existing routers adopt dive…

cs.SE2026

CUADebug: Diagnosing and Repairing Computer-Use Agent Failures

Weijia Zhang, Kunlun Zhu, Zeyi Liu +8

Computer-use agents (CUAs) operate real desktop and web interfaces through screenshots, mouse and keyboard actions, and stateful UI feedback, yet their failures remain difficult to…

cs.AI2026

AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM Agents

Kunlun Zhu, Xuyan Ye, Zhiguang Han +9

LLM agent failures are difficult to debug because the step where an error surfaces is often not the one that caused it. Existing observability tools replay execution traces but pro…

cs.AI2026

BioInsight: Multi-Agent Orchestration for Interactive Biomedical Knowledge Discovery

Jieyi Wang, Bingxuan Li, Nanyi Jiang +9

Biomedical deep-research systems increasingly retrieve and synthesize scientific evidence, but their outputs typically collapse heterogeneous evidence into static text, making prov…

cs.AI2026

ProtocolBench: Which LLM MultiAgent Protocol to Choose?

Hongyi Du, Jiaqi Su, Jisen Li +6

As large-scale multi-agent systems evolve, the communication protocol layer has become a critical yet under-evaluated factor shaping performance and reliability. Despite the existe…