15 papers
Diagnosing the Reliability of LLM-as-a-Judge via Item Response Theory
Junhyuk Choi, Sohhyung Park, Chanhee Cho +2
While LLM-as-a-Judge is widely used in automated evaluation, existing validation practices primarily operate at the level of observed outputs, offering limited insight into whether…
VorTEX: Various overlap ratio for Target speech EXtraction
Ro-hoon Oh, Jihwan Seol, Bugeun Kim
Target speech extraction (TSE) aims to recover a target speaker's voice from a mixture. While recent text-prompted approaches have shown promise, most approaches assume fully overl…
SPAM: Style Prompt Adherence Metric for Prompt-based TTS
Chanhee Cho, Nayeon Kim, Bugeun Kim
Prompt-based text-to-speech (TTS) aims to generate speech that adheres to fine-grained style cues provided in a text prompt. However, most prior works depend on neither plausible n…
Belief in Authority: Impact of Authority in Multi-Agent Evaluation Framework
Junhyuk Choi, Jeongyoun Kwon, Heeju Kim +4
Multi-agent systems utilizing large language models often assign authoritative roles to improve performance, yet the impact of authority bias on agent interactions remains underexp…
Fine-Grained and Thematic Evaluation of LLMs in Social Deduction Game
Byungjun Kim, Dayeon Seo, Minju Kim +1
Recent studies have investigated whether large language models (LLMs) can support obscured communication, which is characterized by core aspects such as inferring subtext and evadi…
Acoustic-based Gender Differentiation in Speech-aware Language Models
Junhyuk Choi, Jihwan Seol, Nayeon Kim +3
Speech-aware Language Models (SpeechLMs) have fundamentally transformed human-AI interaction by enabling voice-based communication, yet they may exhibit acoustic-based gender diffe…