5 papers
KVoiceBench, KOpenAudioBench, and KMMAU: Agent-Driven Korean Speech Benchmarks for Evaluating SpeechLMs
Haechan Kim, Seungjun Chung, Inkyu Park +2
Speech language models (SpeechLMs) have achieved substantial progress by extending large language models (LLMs) to the speech modality. However, SpeechLM evaluation remains heavily…
Looped Diffusion Language Models
Sanghyun Lee, Chunsan Hong, Seungryong Kim +3
Masked diffusion models (MDMs) have emerged as a promising alternative to autoregressive models for language modeling, yet the effective design of transformer architectures for MDM…
MMTB: Evaluating Terminal Agents on Multimedia-File Tasks
Chiyeong Heo, Jaechang Kim, Junhyuk Kwon +4
Terminals provide a powerful interface for AI agents by exposing diverse tools for automating complex workflows, yet existing terminal-agent benchmarks largely focus on tasks groun…
Raon-Speech Technical Report
Beomsoo Kim, Changho Choi, Dohyun Kim +23
We present Raon-Speech, a top-performing 9B-parameter speech language model (SpeechLM) for English and Korean speech understanding, answering, and generation, and Raon-SpeechChat,…
Exploring Multimodal Perception in Large Language Models Through Perceptual Strength Ratings
Jonghyun Lee, Dojun Park, Jiwoo Lee +2
This study investigated whether multimodal large language models can achieve human-like sensory grounding by examining their ability to capture perceptual strength ratings across s…