From the 1 of 2 linked papers with an AI index.
2 papers
cs.CL2026
ORCA-bench: How Ready Are Language Model Agents for Oncall?
Albert Gong, Kyuseong Choi, Abhineet Agarwal +5
The paper presents ORCA-bench, a benchmark that evaluates large language model agents on on-call root cause analysis tasks using real telemetry data from a live microservice system…
cs.IR2026
Atomic Information Flow: A Network Flow Model for Tool Attributions in RAG Systems
James Gao, Josh Zhou, Qi Sun +2
Many tool-based Retrieval Augmented Generation (RAG) systems lack precise mechanisms for tracing final responses back to specific tool components -- a critical gap as systems scale…