7 citations · 7 across the 8 of their papers we have counts for
8 papers
A Composable Evaluation System for Reproducible Omni-Modal Foundation Model Evaluation
Hodong Lee, Sanghee Park, Dohoon Ryu +4
Building an omni-modal foundation model means evaluating it across text, image, video, and audio. Excellent evaluation toolkits exist for each modality, but their inference engines…
Where Does the Sound Go? Tracing Acoustic Information Loss in Audio-Conditioned LLMs
Song-ha Jo, Sehyun Lee, Soyoon Kim +2
Audio-conditioned language models often underuse acoustic cues such as prosody, emotion, and non-speech sounds, raising the question of whether ASR-supervised frontends discard thi…
ParaPairAudioBench: Paralinguistic Pairwise Audio Benchmark for LALM-as-a-Judge
Jisu Jeon, Seungyeon Jwa, Joosung Lee +6
Large Audio-Language Models (LALMs) have been widely used as judge models for the automatic evaluation of generated speech. However, prior approaches predominantly focus on holisti…
Sommelier: Scalable Open Multi-turn Audio Pre-processing for Full-duplex Speech Language Models
Kyudan Jung, Jihwan Kim, Soyoon Kim +3
As the paradigm of AI shifts from text-based LLMs to Speech Language Models (SLMs), there is a growing demand for full-duplex systems capable of real-time, natural human-computer i…
OmniACBench: A Benchmark for Evaluating Context-Grounded Acoustic Control in Omni-Modal Models
Seunghee Kim, Bumkyu Park, Kyudan Jung +5
Most testbeds for omni-modal models assess multimodal understanding via textual outputs, leaving it unclear whether these models can properly speak their answers. To study this, we…
SNAP: Speaker Nulling for Artifact Projection in Speech Deepfake Detection
Kyudan Jung, Jihwan Kim, Minwoo Lee +4
Recent advancements in text-to-speech technologies enable generating high-fidelity synthetic speech nearly indistinguishable from real human voices. While recent studies show the e…