2 papers
cs.SE2026
CodeChat-Eval: Evaluating Large Language Models in Multi-Turn Code Refinement Dialogues
Guoxiang Guo, Kla Tantithamthavorn, Neelofar Neelofar +2
Large Language Models (LLMs) are increasingly used in software engineering to generate and refine code. In practice, developers often continue from an initial code generation reque…
cs.SE2026
MORTAR: Multi-turn Metamorphic Testing for LLM-based Dialogue Systems
Aaron Guoxiang Guo, Aldeida Aleti, Neelofar Neelofar +3
With the widespread application of LLM-based dialogue systems in daily life, quality assurance has become more important than ever. Recent research has successfully introduced meth…