6 papers
Towards Direct Evaluation of Harness Optimizers via Priority Ranking
Kai Tzu-iunn Ong, Minseok Kang, Dongwook Choi +9
Harness optimization enables automated agent creation by having an optimizer agent iteratively update the harness of target agents. Despite its success, current studies evaluate op…
LEGO-Eval: Towards Fine-Grained Evaluation on Synthesizing 3D Embodied Environments with Tool Augmentation
Gyeom Hwangbo, Hyungjoo Chae, Minseok Kang +3
Despite recent progress in using Large Language Models (LLMs) for automatically generating 3D scenes, generated scenes often lack realistic spatial layouts and object attributes fo…
Web-Shepherd: Advancing PRMs for Reinforcing Web Agents
Hyungjoo Chae, Sunghwan Kim, Junhee Cho +18
Web navigation is a unique domain that can automate many repetitive real-life tasks and is challenging as it requires long-horizon sequential decision making beyond typical multimo…
PRINCIPLES: Synthetic Strategy Memory for Proactive Dialogue Agents
Namyoung Kim, Kai Tzu-iunn Ong, Yeonjun Hwang +5
Dialogue agents based on large language models (LLMs) have shown promising performance in proactive dialogue, which requires effective strategy planning. However, existing approach…
Can You Share Your Story? Modeling Clients' Metacognition and Openness for LLM Therapist Evaluation
Minju Kim, Dongje Yoo, Yeonjun Hwang +11
Understanding clients' thoughts and beliefs is fundamental in counseling, yet current evaluations of LLM therapists often fail to assess this ability. Existing evaluation methods r…
Cactus: Towards Psychological Counseling Conversations using Cognitive Behavioral Theory
Suyeon Lee, Sunghwan Kim, Minju Kim +11
Recently, the demand for psychological counseling has significantly increased as more individuals express concerns about their mental health. This surge has accelerated efforts to…