3 papers
cs.AI2026
SentinelBench: A Benchmark for Long-Running Monitoring Agents
Matheus Kunzler Maldaner, Adam Fourney, Amanda Swearngin +5
AI agents are increasingly asked to carry out work that spans minutes, hours, or longer. Yet the default model of agent behavior is continuous action: issuing tool calls, refreshin…
cs.AI2025
Magentic-UI: Towards Human-in-the-loop Agentic Systems
Hussein Mozannar, Gagan Bansal, Cheng Tan +17
AI agents powered by large language models are increasingly capable of autonomously completing complex, multi-step tasks using external tools. Yet, they still fall short of human-l…
cs.AI2024
Magentic-One: A Generalist Multi-Agent System for Solving Complex Tasks
Adam Fourney, Gagan Bansal, Hussein Mozannar +17
Modern AI agents, driven by advances in large foundation models, promise to enhance our productivity and transform our lives by augmenting our knowledge and capabilities. To achiev…