4 papers
Towards Direct Evaluation of Harness Optimizers via Priority Ranking
Kai Tzu-iunn Ong, Minseok Kang, Dongwook Choi +9
Harness optimization enables automated agent creation by having an optimizer agent iteratively update the harness of target agents. Despite its success, current studies evaluate op…
On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon Length
Sunghwan Kim, Junhee Cho, Beong-woo Kwak +6
Large language models (LLMs) have shown promise as interactive agents that solve tasks through extended sequences of environment interactions. While prior work has primarily focuse…
Embodied Agents Meet Personalization: Investigating Challenges and Solutions Through the Lens of Memory Utilization
Taeyoon Kwon, Dongwook Choi, Hyojun Kim +5
LLM-powered embodied agents have shown success on conventional object-rearrangement tasks, but providing personalized assistance that leverages user-specific knowledge from past in…
Stop Playing the Guessing Game! Target-free User Simulation for Evaluating Conversational Recommender Systems
Sunghwan Kim, Kwangwook Seo, Tongyoung Kim +2
Recent approaches in Conversational Recommender Systems (CRSs) have tried to simulate real-world users engaging in conversations with CRSs to create more realistic testing environm…