2 papers
cs.LG2026
STABLEVAL: Disagreement-Aware and Stable Evaluation of AI Systems
Akash Bonagiri, Gerard Janno Anderias, Saee Patil +6
Human evaluation remains the primary standard for assessing modern AI systems, yet annotator disagreement, bias, and variability make system rankings fragile under standard majorit…
cs.LG2026
CausalFlow: Causal Attribution and Counterfactual Repair for LLM Agent Failures
Akash Bonagiri, Devang Borkar, Gerard Janno Anderias +2
Large language model (LLM) agents frequently fail on multi-step tasks involving reasoning, tool use, and environment interaction. While such failures are typically logged or retrie…