4 papers · 1 filter
The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI
Muayad Sayed Ali, Aliaksandra Novik, Anji Boddupally +29
Agentic AI development today runs on token maxing: buying capability with tokens -- longer reasoning traces, more turns, wider tool payloads, bigger replayed contexts -- so tokens…
The Price of Agreement: Measuring LLM Sycophancy in Agentic Financial Applications
Zhenyu Zhao, Aparna Balagopalan, Adi Agrawal +3
Given the increased use of LLMs in financial systems today, it becomes important to evaluate the safety and robustness of such systems. One failure mode that LLMs frequently displa…
Towards Outcome-Oriented, Task-Agnostic Evaluation of AI Agents
Waseem AlShikh, Muayad Sayed Ali, Brian Kennedy +1
As AI agents proliferate across industries and applications, evaluating their performance based solely on infrastructural metrics such as latency, time-to-first-token, or token thr…
OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web
Raghav Kapoor, Yash Parag Butala, Melisa Russak +4
For decades, human-computer interaction has fundamentally been manual. Even today, almost all productive work done on the computer necessitates human input at every step. Autonomou…