Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
MDSEval: A Meta-Evaluation Benchmark for Multimodal Dialogue Summarization
Yinhong Liu, Jianfeng He, Hang Su +6
Multimodal Dialogue Summarization (MDS) is a critical task with wide-ranging applications. To support the development of effective MDS models, robust automatic evaluation methods a…
cs.CL2025
Peacemaker or Troublemaker: How Sycophancy Shapes Multi-Agent Debate
Binwei Yao, Chao Shang, Wanyu Du +6
Large language models (LLMs) often display sycophancy, a tendency toward excessive agreeability. This behavior poses significant challenges for multi-agent debating systems (MADS)…