4 citations · 5 across the 3 of their papers we have counts for
3 papers
MEMTRACK: Evaluating Long-Term Memory and State Tracking in Multi-Platform Dynamic Agent Environments
Darshan Deshpande, Varun Gangal, Hersh Mehta +3
Recent works on context and memory benchmarking have primarily focused on conversational instances but the need for evaluating memory in dynamic enterprise environments is crucial…
GLIDER: Grading LLM Interactions and Decisions using Explainable Ranking
Darshan Deshpande, Selvan Sunitha Ravi, Sky CH-Wang +3
The LLM-as-judge paradigm is increasingly being adopted for automated evaluation of model outputs. While LLM judges have shown promise on constrained evaluation tasks, closed sourc…
Lynx: An Open Source Hallucination Evaluation Model
Selvan Sunitha Ravi, Bartosz Mielczarek, Anand Kannappan +2
Retrieval Augmented Generation (RAG) techniques aim to mitigate hallucinations in Large Language Models (LLMs). However, LLMs can still produce information that is unsupported or c…