Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
DETOUR: An Interactive Benchmark for Dual-Agent Search and Reasoning
Li Siyan, Darshan Deshpande, Anand Kannappan +1
When recalling information in conversation, people often arrive at the recollection after multiple turns. However, existing benchmarks for evaluating agent capabilities in such tip…
cs.CL2024
GLIDER: Grading LLM Interactions and Decisions using Explainable Ranking
Darshan Deshpande, Selvan Sunitha Ravi, Sky CH-Wang +3
The LLM-as-judge paradigm is increasingly being adopted for automated evaluation of model outputs. While LLM judges have shown promise on constrained evaluation tasks, closed sourc…
cs.CL2024
GNOME: Generating Negotiations through Open-Domain Mapping of Exchanges
Darshan Deshpande, Shambhavi Sinha, Anirudh Ravi Kumar +2
Language Models have previously shown strong negotiation capabilities in closed domains where the negotiation strategy prediction scope is constrained to a specific setup. In this…