10 papers
Towards Direct Evaluation of Harness Optimizers via Priority Ranking
Kai Tzu-iunn Ong, Minseok Kang, Dongwook Choi +9
Harness optimization enables automated agent creation by having an optimizer agent iteratively update the harness of target agents. Despite its success, current studies evaluate op…
PRINCIPLES: Synthetic Strategy Memory for Proactive Dialogue Agents
Namyoung Kim, Kai Tzu-iunn Ong, Yeonjun Hwang +5
Dialogue agents based on large language models (LLMs) have shown promising performance in proactive dialogue, which requires effective strategy planning. However, existing approach…
Web Agents with World Models: Learning and Leveraging Environment Dynamics in Web Navigation
Hyungjoo Chae, Namyoung Kim, Kai Tzu-iunn Ong +6
Large language models (LLMs) have recently gained much attention in building autonomous agents. However, the performance of current LLM-based web agents in long-horizon tasks is fa…
Towards Lifelong Dialogue Agents via Timeline-based Memory Management
Kai Tzu-iunn Ong, Namyoung Kim, Minju Gwak +6
To achieve lifelong human-agent interaction, dialogue agents need to constantly memorize perceived information and properly retrieve it for response generation (RG). While prior st…
Coffee-Gym: An Environment for Evaluating and Improving Natural Language Feedback on Erroneous Code
Hyungjoo Chae, Taeyoon Kwon, Seungjun Moon +7
This paper presents Coffee-Gym, a comprehensive RL environment for training models that provide feedback on code editing. Coffee-Gym includes two major components: (1) Coffee, a da…
Large Language Models Are Self-Taught Reasoners: Enhancing LLM Applications via Tailored Problem-Solving Demonstrations
Kai Tzu-iunn Ong, Taeyoon Kwon, Jinyoung Yeo
Guiding large language models with a selected set of human-authored demonstrations is a common practice for improving LLM applications. However, human effort can be costly, especia…