activity
20242026
collaborators

5 papers

cs.CL2026

Accurate Failure Prediction in Agents Does Not Imply Effective Failure Prevention

Rakshith Vasudev, Melisa Russak, Dan Bikel +1

Proactive interventions by LLM critic models are often assumed to improve reliability, yet their effects at deployment time are poorly understood. We show that a binary LLM critic…

cs.AI2025

Towards Outcome-Oriented, Task-Agnostic Evaluation of AI Agents

Waseem AlShikh, Muayad Sayed Ali, Brian Kennedy +1

As AI agents proliferate across industries and applications, evaluating their performance based solely on infrastructural metrics such as latency, time-to-first-token, or token thr…

cs.CL2025

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning

Shelly Bensal, Umar Jamil, Christopher Bryant +5

We explore a method for improving the performance of large language models through self-reflection and reinforcement learning. By incentivizing the model to generate better self-re…

cs.CL2025

Expect the Unexpected: FailSafe Long Context QA for Finance

Kiran Kamble, Melisa Russak, Dmytro Mozolevskyi +3

We propose a new long-context financial benchmark, FailSafeQA, designed to test the robustness and context-awareness of LLMs against six variations in human-interface interactions…

cs.IR2024

Comparative Analysis of Retrieval Systems in the Real World

Dmytro Mozolevskyi, Waseem AlShikh

This research paper presents a comprehensive analysis of integrating advanced language models with search and retrieval systems in the fields of information retrieval and natural l…