32 papers · 1 filter
ArcANE: Do Role-Playing Language Agents Stay in Character at the Right Time?
Woojung Song, Nalim Kim, Sangjun Song +3
Role-playing language agents (RPLAs) should play characters whose values and behavior evolve as the story progresses, not maintain a fixed persona. Existing benchmarks measure fact…
Bridging the Knowledge-Prediction Gap in LLMs on Multiple-Choice Questions
Yoonah Park, Haesung Pyun, Yohan Jo
While large language models (LLMs) perform strongly on diverse tasks, their trustworthiness is limited by erratic behavior that is unfaithful to their internal knowledge. In partic…
Dual Mechanisms of Value Expression: Intrinsic vs. Prompted Values in Large Language Models
Jongwook Han, Jongwon Lim, Injin Kong +1
Large language models can express values in two main ways: (1) intrinsic expression, reflecting the model's inherent values learned during training, and (2) prompted expression, el…
Human Psychometric Questionnaires Mischaracterize LLM Behavior
Woojung Song, Dongmin Choi, Yoonah Park +3
We examine whether human psychometric questionnaires can serve as reliable tools for characterizing and predicting LLM behavior in everyday user interactions. We analyze eight open…
Psychometric Item Validation Using Virtual Respondents with Trait-Response Mediators
Sungjib Lim, Woojung Song, Eun-Ju Lee +1
As psychometric surveys are increasingly used to assess the traits of large language models (LLMs), the need for scalable survey item generation suited for LLMs has also grown. A c…
Where Should Diffusion Enter a Language Model? Geometry-Guided Hidden-State Replacement
Injin Kong, Hyoungjoon Lee, Yohan Jo
Continuous diffusion language models lag behind autoregressive transformers, partly because diffusion is applied in spaces poorly suited to language denoising and token recovery. W…