3 papers
cs.SE2026
AgentLens: Revealing The Lucky Pass Problem in SWE-Agent Evaluation
Priyam Sahoo, Gaurav Mittal, Xiaomin Li +4
Evaluation of software engineering (SWE) agents is dominated by a binary signal: whether the final patch passes the tests. This outcome-only view treats a principled solution and a…
cs.CL2026
DialDefer: A Framework for Detecting and Mitigating LLM Dialogic Deference
Parisa Rabbani, Priyam Sahoo, Ruben Mathew +4
LLMs are increasingly used as third-party judges, yet their reliability when evaluating speakers in dialogue remains poorly understood. We show that LLMs judge identical claims dif…
cs.SE2024
Insights from the Usage of the Ansible Lightspeed Code Completion Service
Priyam Sahoo, Saurabh Pujar, Ganesh Nalawade +3
The availability of Large Language Models (LLMs) which can generate code, has made it possible to create tools that improve developer productivity. Integrated development environme…