1 paper · 1 filter
Ali Pourghasemi Fatideh, Wilder Baldwin, Maria Dhakal +2
LLM-based dialogue assistants have become mainstream tools for software developers, yet current evaluation benchmarks focus exclusively on functional correctness. This leaves a cri…