From the 1 of 17 linked papers with an AI index.
1 paper · 1 filter
Ali Pourghasemi Fatideh, Wilder Baldwin, Maria Dhakal +2
LLM-based dialogue assistants have become mainstream tools for software developers, yet current evaluation benchmarks focus exclusively on functional correctness. This leaves a cri…