1 citations · 1 across the 4 of their papers we have counts for
4 papers
DETOUR: An Interactive Benchmark for Dual-Agent Search and Reasoning
Li Siyan, Darshan Deshpande, Anand Kannappan +1
When recalling information in conversation, people often arrive at the recollection after multiple turns. However, existing benchmarks for evaluating agent capabilities in such tip…
Benchmarking Reward Hack Detection in Code Environments via Contrastive Analysis
Darshan Deshpande, Anand Kannappan, Rebecca Qian
Recent advances in reinforcement learning for code generation have made robust environments essential to prevent reward hacking. As LLMs increasingly serve as evaluators in code-ba…
MEMTRACK: Evaluating Long-Term Memory and State Tracking in Multi-Platform Dynamic Agent Environments
Darshan Deshpande, Varun Gangal, Hersh Mehta +3
Recent works on context and memory benchmarking have primarily focused on conversational instances but the need for evaluating memory in dynamic enterprise environments is crucial…
GLIDER: Grading LLM Interactions and Decisions using Explainable Ranking
Darshan Deshpande, Selvan Sunitha Ravi, Sky CH-Wang +3
The LLM-as-judge paradigm is increasingly being adopted for automated evaluation of model outputs. While LLM judges have shown promise on constrained evaluation tasks, closed sourc…