#ai agents
11 papers match
Stop Shipping AI Agents on Faith: Capability Is Not Production Readiness
Fouad Bousetouane
The paper proposes the ProofAgent Index (PAI), a governance readiness framework for AI agents that assesses evaluation, context, compliance, and governance to determine production…
Open Security Benchmark: Towards Autonomous Enterprise Cyber Defense
Gal Engelberg, Michael Arenzon, Leon Goldberg
The paper introduces the Open Security Benchmark (OSB), a framework that provides a frozen, holistic enterprise security dataset and evaluation tools for testing autonomous AI agen…
Can AI agents conduct open-ended AI research? Early evidence from two case studies
Peter Kirgis, Sayash Kapoor, Andrew Schwartz +21
The paper evaluates whether current AI agents can independently conduct open‑ended AI research by having them attempt to solve the central questions of two unpublished NeurIPS subm…
What Does It Take to Detect an AI Agent? Minimal Feature Sets for Behavioral Detection under Browser Automation
Vishisht Choudhary, Lukas Schmidt, Anne Zoë Kenntner +3
The paper introduces a three‑class detection framework that distinguishes humans, bots, and AI agents browsing via automation, showing that a few behavioral features can reliably i…
Toward Standardized Cross-Vendor Agent Tool Trust Management in Autonomous Networks
Ravi Kant Sharma, Ashutosh Uttam, Ajay Kumar
The paper proposes AgentToolMO, a 3GPP NRM information model that enables AI agents in autonomous networks to manage and share trust information about cross‑vendor tools, limiting…
LQCDMaster: Agentic Scientific Computing for Lattice Quantum Chromodynamics Research
Haofei Gao, Tingjia Miao, Wenkai Jin +12
LQCDMaster is an AI-driven scientific computing agent that translates natural‑language lattice QCD research tasks into fully executable PyQUDA workflows, automating code generation…