38 papers
ArcANE: Do Role-Playing Language Agents Stay in Character at the Right Time?
Woojung Song, Nalim Kim, Sangjun Song +3
Role-playing language agents (RPLAs) should play characters whose values and behavior evolve as the story progresses, not maintain a fixed persona. Existing benchmarks measure fact…
RobotValues: Evaluating Household Robots When Human Values Conflict
Jongwook Han, Hyeongjin Kim, Yohan Jo
While household robots are often evaluated based on task completion, everyday domestic environments involve value-conflicting situations in which robots are expected to choose acti…
Bridging the Knowledge-Prediction Gap in LLMs on Multiple-Choice Questions
Yoonah Park, Haesung Pyun, Yohan Jo
While large language models (LLMs) perform strongly on diverse tasks, their trustworthiness is limited by erratic behavior that is unfaithful to their internal knowledge. In partic…
Dual Mechanisms of Value Expression: Intrinsic vs. Prompted Values in Large Language Models
Jongwook Han, Jongwon Lim, Injin Kong +1
Large language models can express values in two main ways: (1) intrinsic expression, reflecting the model's inherent values learned during training, and (2) prompted expression, el…
Human Psychometric Questionnaires Mischaracterize LLM Behavior
Woojung Song, Dongmin Choi, Yoonah Park +3
We examine whether human psychometric questionnaires can serve as reliable tools for characterizing and predicting LLM behavior in everyday user interactions. We analyze eight open…
Mechanism Shift During Post-training from Autoregressive to Masked Diffusion Language Models
Injin Kong, Hyoungjoon Lee, Yohan Jo
Post-training pretrained autoregressive models (ARMs) into masked diffusion models (MDMs) has emerged as a cost-effective way to overcome the limitations of sequential generation.…