Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures
Harsh Raj, Vipul Gupta, Anas Mahmoud +4
Existing evaluations often reduce agent failures to system-level outcomes, obscuring where the fault originated and which intervention would improve the agent system. This creates…
cs.AI2026
OpenThoughts-Agent: Data Recipes for Agentic Models
Negin Raoof, Richard Zhuang, Marianna Nezhurina +47
Agentic language models dramatically expand the applications of AI yet little is publicly known about how to curate training data for broadly capable agents. Existing open efforts…
cs.AI2026
Consistency as a Testable Property: Statistical Methods to Evaluate AI Agent Reliability
Harsh Raj, Niranjan Orkat, Suvrorup Mukherjee +3
This paper establishes a rigorous measurement science for AI agent reliability, providing a foundational framework for quantifying consistency under semantically preserving perturb…