6 papers
Emergence World: A Platform for Evaluating Long-Horizon Multi-Agent Autonomy
Deepak Akkil, Ravi Kokku, Karthik Vikram +3
Most evaluations of LLM agents look like exams: a discrete task, a clean environment, a score in minutes or hours. We argue that this approach is mismatched with the deployment con…
Learning API Functionality from In-Context Demonstrations for Tool-based Agents
Bhrij Patel, Ashish Jagmohan, Aditya Vempaty
Digital tool-based agents, powered by Large Language Models (LLMs), that invoke external Application Programming Interfaces (APIs) often rely on documentation to understand API fun…
Introducing Spotlight: A Novel Approach for Generating Captivating Key Information from Documents
Ankan Mullick, Sombit Bose, Rounak Saha +6
In this paper, we introduce Spotlight, a novel paradigm for information extraction that produces concise, engaging narratives by highlighting the most compelling aspects of a docum…
Reflection-Based Memory For Web navigation Agents
Ruhana Azam, Aditya Vempaty, Ashish Jagmohan
Web navigation agents have made significant progress, yet current systems operate with no memory of past experiences -- leading to repeated mistakes and an inability to learn from…
Better RAG using Relevant Information Gain
Marc Pickett, Jeremy Hartman, Ayan Kumar Bhowmick +2
A common way to extend the memory of large language models (LLMs) is by retrieval augmented generation (RAG), which inserts text retrieved from a larger memory into an LLM's contex…
Multimodal Auto Validation For Self-Refinement in Web Agents
Ruhana Azam, Tamer Abuelsaad, Aditya Vempaty +1
As our world digitizes, web agents that can automate complex and monotonous tasks are becoming essential in streamlining workflows. This paper introduces an approach to improving w…