activity
20232026
most citedUser Simulation with Large Language Models for Evaluating Task-Oriented Dialogue

4 citations · 11 across the 8 of their papers we have counts for

collaborators
Showing cs.CLShow all

7 papers · 1 filter

cs.CL20242 cited

FineSurE: Fine-grained Summarization Evaluation using LLMs

Hwanjun Song, Hang Su, Igor Shalyminov +2

Automated evaluation is crucial for streamlining text summarization benchmarking and model development, given the costly and time-consuming nature of human evaluation. Traditional…

cs.CL20242 cited

TofuEval: Evaluating Hallucinations of LLMs on Topic-Focused Dialogue Summarization

Liyan Tang, Igor Shalyminov, Amy Wing-mei Wong +11

Single document news summarization has seen substantial progress on faithfulness in recent years, driven by research on the evaluation of factual consistency, or hallucinations. We…

cs.CL2024

Can Your Model Tell a Negation from an Implicature? Unravelling Challenges With Intent Encoders

Yuwei Zhang, Siffi Singh, Sailik Sengupta +4

Conversational systems often rely on embedding models for intent classification and intent clustering tasks. The advent of Large Language Models (LLMs), which enable instructional…

cs.CL2024

Semi-Supervised Dialogue Abstractive Summarization via High-Quality Pseudolabel Selection

Jianfeng He, Hang Su, Jason Cai +3

Semi-supervised dialogue summarization (SSDS) leverages model-generated summaries to reduce reliance on human-labeled data and improve the performance of summarization models. Whil…

cs.CL20234 cited

User Simulation with Large Language Models for Evaluating Task-Oriented Dialogue

Sam Davidson, Salvatore Romeo, Raphael Shu +4

One of the major impediments to the development of new task-oriented dialogue (TOD) systems is the need for human evaluation at multiple stages and iterations of the development pr…

cs.CL2023

NatCS: Eliciting Natural Customer Support Dialogues

James Gung, Emily Moeng, Wesley Rose +3

Despite growing interest in applications based on natural customer support conversations, there exist remarkably few publicly available datasets that reflect the expected character…