activity
20232026
most citedUser Simulation with Large Language Models for Evaluating Task-Oriented Dialogue

4 citations · 9 across the 13 of their papers we have counts for

collaborators
Showing cs.CLShow all

10 papers · 1 filter

cs.CL2026

ConvApparel: A Benchmark Dataset and Validation Framework for User Simulators in Conversational Recommenders

Ofer Meshi, Krisztian Balog, Sally Goldman +5

The promise of LLM-based user simulators to improve conversational AI is hindered by a critical "realism gap," leading to systems that are optimized for simulated interactions, but…

cs.CL2025

Controllable Conversational Theme Detection Track at DSTC 12

Igor Shalyminov, Hang Su, Jake Vincent +5

Conversational analytics has been on the forefront of transformation driven by the advances in Speech and Natural Language Processing techniques. Rapid adoption of Large Language M…

cs.CL2025

CONFETTI: Conversational Function-Calling Evaluation Through Turn-Level Interactions

Tamer Alkhouli, Katerina Margatina, James Gung +4

We introduce Conversational Function-Calling Evaluation Through Turn-Level Interactions (CONFETTI), a conversational benchmark1 designed to evaluate the function-calling capabiliti…

cs.CL20241 cited

Structured List-Grounded Question Answering

Mujeen Sung, Song Feng, James Gung +3

Document-grounded dialogue systems aim to answer user queries by leveraging external information. Previous studies have mainly focused on handling free-form documents, often overlo…

cs.CL2024

DFlow: Diverse Dialogue Flow Simulation with Large Language Models

Wanyu Du, Song Feng, James Gung +4

Developing language model-based dialogue agents requires effective data to train models that can follow specific task logic. However, most existing data simulation methods focus on…

cs.CL20234 cited

User Simulation with Large Language Models for Evaluating Task-Oriented Dialogue

Sam Davidson, Salvatore Romeo, Raphael Shu +4

One of the major impediments to the development of new task-oriented dialogue (TOD) systems is the need for human evaluation at multiple stages and iterations of the development pr…