8 papers
Agents Trust Tools Too Much: Measuring Reliance on Unreliable Tools
Hoyeol Yang, Woojung Song, Taewon Kim +3
Existing evaluations of tool-using agents primarily measure whether an agent can successfully complete diverse tasks with tools. These evaluations generally assume that tools retur…
ArcANE: Do Role-Playing Language Agents Stay in Character at the Right Time?
Woojung Song, Nalim Kim, Sangjun Song +3
Role-playing language agents (RPLAs) simulate specific characters and personas across applications such as entertainment, companionship, interactive storytelling, and education. Fa…
Don't Adapt Small Language Models for Tools; Adapt Tool Schemas to the Models
Jonggeun Lee, Woojung Song, Jongwook Han +2
Small language models (SLMs) enable scalable tool-augmented multi-agent systems where multiple SLMs handle subtasks orchestrated by a powerful coordinator. However, they struggle w…
Quantifying Data Contamination in Psychometric Evaluations of LLMs
Jongwook Han, Woojung Song, Jonggeun Lee +1
Recent studies apply psychometric questionnaires to Large Language Models (LLMs) to assess high-level psychological constructs such as values, personality, moral foundations, and d…
Non-Collaborative User Simulators for Tool Agents
Jeonghoon Shim, Woojung Song, Cheyon Jin +2
Tool agents interact with users through multi-turn dialogues to accomplish various tasks. Recent studies have adopted user simulation methods to develop these agents in multi-turn…
Human Psychometric Questionnaires Mischaracterize LLM Behavior
Woojung Song, Dongmin Choi, Yoonah Park +3
We examine whether human psychometric questionnaires can serve as reliable tools for characterizing and predicting LLM behavior in everyday user interactions. We analyze eight open…