collaborators

11 papers

cs.AI2026

Diagnosing the Reliability of LLM-as-a-Judge via Item Response Theory

Junhyuk Choi, Sohhyung Park, Chanhee Cho +2

While LLM-as-a-Judge is widely used in automated evaluation, existing validation practices primarily operate at the level of observed outputs, offering limited insight into whether…

cs.CL2026

Framing Matters: Addressing Framing Sensitivity in Decision-Making through Behaviorally-Grounded Value Alignment

Seojin Hwang, Minju Kim, Junhyuk Choi +2

Large Language Models (LLMs) are increasingly deployed in high-stakes decision-making settings such as legal reasoning, where consistency under factually equivalent inputs is criti…

cs.CL2026

Identifying and Mitigating Bottlenecks in Role-Playing Agents: A Systematic Study of Disentangling Character Profile Axes

Yonghyun Jun, Junhyuk Choi, Jeonghyun Park +3

While Large Language Model (LLM) role-playing agents have advanced rapidly, it remains unclear which profile elements genuinely drive role-playing quality. To bridge this gap, we i…

cs.CL2026

Belief in Authority: Impact of Authority in Multi-Agent Evaluation Framework

Junhyuk Choi, Jeongyoun Kwon, Heeju Kim +4

Multi-agent systems utilizing large language models often assign authoritative roles to improve performance, yet the impact of authority bias on agent interactions remains underexp…

cs.CL2025

Acoustic-based Gender Differentiation in Speech-aware Language Models

Junhyuk Choi, Jihwan Seol, Nayeon Kim +3

Speech-aware Language Models (SpeechLMs) have fundamentally transformed human-AI interaction by enabling voice-based communication, yet they may exhibit acoustic-based gender diffe…

cs.CL2025

VoiceBBQ: Investigating Effect of Content and Acoustics in Social Bias of Spoken Language Model

Junhyuk Choi, Ro-hoon Oh, Jihwan Seol +1

We introduce VoiceBBQ, a spoken extension of the BBQ (Bias Benchmark for Question Answering) - a dataset that measures social bias by presenting ambiguous or disambiguated contexts…