1 citations · 1 across the 5 of their papers we have counts for
6 papers
PAC-BENCH: Evaluating Multi-Agent Collaboration under Privacy Constraints
Minjun Park, Donghyun Kim, Hyeonjong Ju +5
We are entering an era in which individuals and organizations increasingly deploy dedicated AI agents that interact and collaborate with other agents. However, the dynamics of mult…
CONDESION-BENCH: Conditional Decision-Making of Large Language Models in Compositional Action Space
Yeonjun Hwang, Sungyong Park, Minju Kim +2
Large language models have been widely explored as decision-support tools in high-stakes domains due to their contextual understanding and reasoning capabilities. However, existing…
Can You Share Your Story? Modeling Clients' Metacognition and Openness for LLM Therapist Evaluation
Minju Kim, Dongje Yoo, Yeonjun Hwang +11
Understanding clients' thoughts and beliefs is fundamental in counseling, yet current evaluations of LLM therapists often fail to assess this ability. Existing evaluation methods r…
ToolHaystack: Stress-Testing Tool-Augmented Language Models in Realistic Long-Term Interactions
Beong-woo Kwak, Minju Kim, Dongha Lim +5
Large language models (LLMs) have demonstrated strong capabilities in using external tools to address user inquiries. However, most existing evaluations assume tool use in short co…
BotsTalk: Machine-sourced Framework for Automatic Curation of Large-scale Multi-skill Dialogue Datasets
Minju Kim, Chaehyeong Kim, Yongho Song +2
To build open-domain chatbots that are able to use diverse communicative skills, we propose a novel framework BotsTalk, where multiple agents grounded to the specific target skills…
Dual Task Framework for Improving Persona-grounded Dialogue Dataset
Minju Kim, Beong-woo Kwak, Youngwook Kim +3
This paper introduces a simple yet effective data-centric approach for the task of improving persona-conditioned dialogue agents. Prior model-centric approaches unquestioningly dep…