11 papers
Diagnosing the Reliability of LLM-as-a-Judge via Item Response Theory
Junhyuk Choi, Sohhyung Park, Chanhee Cho +2
While LLM-as-a-Judge is widely used in automated evaluation, existing validation practices primarily operate at the level of observed outputs, offering limited insight into whether…
Framing Matters: Addressing Framing Sensitivity in Decision-Making through Behaviorally-Grounded Value Alignment
Seojin Hwang, Minju Kim, Junhyuk Choi +2
Large Language Models (LLMs) are increasingly deployed in high-stakes decision-making settings such as legal reasoning, where consistency under factually equivalent inputs is criti…
Identifying and Mitigating Bottlenecks in Role-Playing Agents: A Systematic Study of Disentangling Character Profile Axes
Yonghyun Jun, Junhyuk Choi, Jeonghyun Park +3
While Large Language Model (LLM) role-playing agents have advanced rapidly, it remains unclear which profile elements genuinely drive role-playing quality. To bridge this gap, we i…
Belief in Authority: Impact of Authority in Multi-Agent Evaluation Framework
Junhyuk Choi, Jeongyoun Kwon, Heeju Kim +4
Multi-agent systems utilizing large language models often assign authoritative roles to improve performance, yet the impact of authority bias on agent interactions remains underexp…
Acoustic-based Gender Differentiation in Speech-aware Language Models
Junhyuk Choi, Jihwan Seol, Nayeon Kim +3
Speech-aware Language Models (SpeechLMs) have fundamentally transformed human-AI interaction by enabling voice-based communication, yet they may exhibit acoustic-based gender diffe…
VoiceBBQ: Investigating Effect of Content and Acoustics in Social Bias of Spoken Language Model
Junhyuk Choi, Ro-hoon Oh, Jihwan Seol +1
We introduce VoiceBBQ, a spoken extension of the BBQ (Bias Benchmark for Question Answering) - a dataset that measures social bias by presenting ambiguous or disambiguated contexts…