1 paper
Ali Pourghasemi Fatideh, Wilder Baldwin, Maria Dhakal +2
LLM-based dialogue assistants have become mainstream tools for software developers, yet current evaluation benchmarks focus exclusively on functional correctness. This leaves a cri…