178 citations · 200 across the 8 of their papers we have counts for
5 papers · 1 filter
SentinelBench: A Benchmark for Long-Running Monitoring Agents
Matheus Kunzler Maldaner, Adam Fourney, Amanda Swearngin +5
AI agents are increasingly asked to carry out work that spans minutes, hours, or longer. Yet the default model of agent behavior is continuous action: issuing tool calls, refreshin…
Magentic-UI: Towards Human-in-the-loop Agentic Systems
Hussein Mozannar, Gagan Bansal, Cheng Tan +17
AI agents powered by large language models are increasingly capable of autonomously completing complex, multi-step tasks using external tools. Yet, they still fall short of human-l…
Measuring AI agent autonomy: Towards a scalable approach with code inspection
Peter Cihon, Merlin Stein, Gagan Bansal +2
AI agents are AI systems that can achieve complex goals autonomously. Assessing the level of agent autonomy is crucial for understanding both their potential benefits and risks. Cu…
Magentic-One: A Generalist Multi-Agent System for Solving Complex Tasks
Adam Fourney, Gagan Bansal, Hussein Mozannar +17
Modern AI agents, driven by advances in large foundation models, promise to enhance our productivity and transform our lives by augmenting our knowledge and capabilities. To achiev…
AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation
Qingyun Wu, Gagan Bansal, Jieyu Zhang +11
AutoGen is an open-source framework that allows developers to build LLM applications via multiple agents that can converse with each other to accomplish tasks. AutoGen agents are c…