most citedOdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

1 citations · 1 across the 7 of their papers we have counts for

collaborators

18 papers

cs.SE2025

Continuous Benchmark Generation for Evaluating Enterprise-scale LLM Agents

Divyanshu Saxena, Rishikesh Maurya, Xiaoxuan Ou +7

The rapid adoption of AI agents across domains has made systematic evaluation crucial for ensuring their usefulness and successful production deployment. Evaluation of AI agents ty…

cs.AI2025

Simulating Environments with Reasoning Models for Agent Training

Yuetai Li, Huseyin A Inan, Xiang Yue +6

LLM agents excel in compact environments requiring deep reasoning but remain brittle when operating in broader, more complex contexts that demand robustness across diverse tools an…

cs.CL2025

Cost-Aware Retrieval-Augmentation Reasoning Models with Adaptive Retrieval Depth

Helia Hashemi, Victor Rühle, Saravan Rajmohan

Reasoning models have gained significant attention due to their strong performance, particularly when enhanced with retrieval augmentation. However, these models often incur high c…

cs.AI2025

LEGOMem: Modular Procedural Memory for Multi-agent LLM Systems for Workflow Automation

Dongge Han, Camille Couturier, Daniel Madrigal Diaz +3

We introduce LEGOMem, a modular procedural memory framework for multi-agent large language model (LLM) systems in workflow automation. LEGOMem decomposes past task trajectories int…

cs.CL20251 cited

OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Weixuan Wang, Dongge Han, Daniel Madrigal Diaz +3

Autonomous agents powered by large language models (LLMs) are increasingly deployed in real-world applications requiring complex, long-horizon workflows. However, existing benchmar…

cs.LG2025

Enhancing Reasoning Capabilities of Small Language Models with Blueprints and Prompt Template Search

Dongge Han, Menglin Xia, Daniel Madrigal Diaz +7

Small language models (SLMs) offer promising and efficient alternatives to large language models (LLMs). However, SLMs' limited capacity restricts their reasoning capabilities and…