7 papers
The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI
Muayad Sayed Ali, Aliaksandra Novik, Anji Boddupally +29
Agentic AI development today runs on token maxing: buying capability with tokens -- longer reasoning traces, more turns, wider tool payloads, bigger replayed contexts -- so tokens…
The Price of Agreement: Measuring LLM Sycophancy in Agentic Financial Applications
Zhenyu Zhao, Aparna Balagopalan, Adi Agrawal +3
Given the increased use of LLMs in financial systems today, it becomes important to evaluate the safety and robustness of such systems. One failure mode that LLMs frequently displa…
Shorthand for Thought: Compressing LLM Reasoning via Entropy-Guided Supertokens
Zhenyu Zhao, Sander Land, Daniel M. Bikel +1
Reasoning in Large Language Models incurs significant inference-time compute, yet the token-level information structure of reasoning traces remains underexplored. We observe that r…
Accurate Failure Prediction in Agents Does Not Imply Effective Failure Prevention
Rakshith Vasudev, Melisa Russak, Dan Bikel +1
Proactive interventions by LLM critic models are often assumed to improve reliability, yet their effects at deployment time are poorly understood. We show that a binary LLM critic…
Towards Outcome-Oriented, Task-Agnostic Evaluation of AI Agents
Waseem AlShikh, Muayad Sayed Ali, Brian Kennedy +1
As AI agents proliferate across industries and applications, evaluating their performance based solely on infrastructural metrics such as latency, time-to-first-token, or token thr…
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning
Shelly Bensal, Umar Jamil, Christopher Bryant +5
We explore a method for improving the performance of large language models through self-reflection and reinforcement learning. By incentivizing the model to generate better self-re…