#ai agents

try —

11 papers match

cs.MA2026

Stop Shipping AI Agents on Faith: Capability Is Not Production Readiness

Fouad Bousetouane

The paper proposes the ProofAgent Index (PAI), a governance readiness framework for AI agents that assesses evaluation, context, compliance, and governance to determine production…

#ai agents#production readiness#governance#evaluation
cs.CR2026

Open Security Benchmark: Towards Autonomous Enterprise Cyber Defense

Gal Engelberg, Michael Arenzon, Leon Goldberg

The paper introduces the Open Security Benchmark (OSB), a framework that provides a frozen, holistic enterprise security dataset and evaluation tools for testing autonomous AI agen…

#autonomous cyber defense#security posture management#benchmarking#enterprise security
cs.AI2026

Can AI agents conduct open-ended AI research? Early evidence from two case studies

Peter Kirgis, Sayash Kapoor, Andrew Schwartz +21

The paper evaluates whether current AI agents can independently conduct open‑ended AI research by having them attempt to solve the central questions of two unpublished NeurIPS subm…

#ai agents#research automation#evaluation methodology#failure analysis
cs.AI2026

What Does It Take to Detect an AI Agent? Minimal Feature Sets for Behavioral Detection under Browser Automation

Vishisht Choudhary, Lukas Schmidt, Anne Zoë Kenntner +3

The paper introduces a three‑class detection framework that distinguishes humans, bots, and AI agents browsing via automation, showing that a few behavioral features can reliably i…

#bot detection#browser automation#AI agents#behavioral features
cs.AI2026

Toward Standardized Cross-Vendor Agent Tool Trust Management in Autonomous Networks

Ravi Kant Sharma, Ashutosh Uttam, Ajay Kumar

The paper proposes AgentToolMO, a 3GPP NRM information model that enables AI agents in autonomous networks to manage and share trust information about cross‑vendor tools, limiting…

#cross-vendor trust#autonomous networks#AI agents#3gpp nrm
hep-lat2026

LQCDMaster: Agentic Scientific Computing for Lattice Quantum Chromodynamics Research

Haofei Gao, Tingjia Miao, Wenkai Jin +12

LQCDMaster is an AI-driven scientific computing agent that translates natural‑language lattice QCD research tasks into fully executable PyQUDA workflows, automating code generation…

#lattice qcd#scientific computing#AI agents#code generation