collaborators

5 papers

cs.CL2026

GLEAN: Active Generalized Category Discovery with Diverse LLM Feedback

Henry Peng Zou, Siffi Singh, Yi Nian +4

Generalized Category Discovery (GCD) is a practical and challenging open-world task that aims to recognize both known and novel categories in unlabeled data using limited labeled d…

cs.CR2026

STAC: When Innocent Tools Form Dangerous Chains for LLM Agents

Jing-Jing Li, Jianfeng He, Chao Shang +6

As LLMs advance into autonomous agents with tool-use capabilities, they introduce security challenges that extend beyond traditional content-based LLM safety concerns. This paper i…

cs.CL2025

MDSEval: A Meta-Evaluation Benchmark for Multimodal Dialogue Summarization

Yinhong Liu, Jianfeng He, Hang Su +6

Multimodal Dialogue Summarization (MDS) is a critical task with wide-ranging applications. To support the development of effective MDS models, robust automatic evaluation methods a…

cs.CL2025

Peacemaker or Troublemaker: How Sycophancy Shapes Multi-Agent Debate

Binwei Yao, Chao Shang, Wanyu Du +6

Large language models (LLMs) often display sycophancy, a tendency toward excessive agreeability. This behavior poses significant challenges for multi-agent debating systems (MADS)…

cs.CL2025

Faithful, Unfaithful or Ambiguous? Multi-Agent Debate with Initial Stance for Summary Evaluation

Mahnaz Koupaee, Jake W. Vincent, Saab Mansour +9

Faithfulness evaluators based on large language models (LLMs) are often fooled by the fluency of the text and struggle with identifying errors in the summaries. We propose an appro…