5 papers
Looped Diffusion Language Models
Sanghyun Lee, Chunsan Hong, Seungryong Kim +3
Masked diffusion models (MDMs) have emerged as a promising alternative to autoregressive models for language modeling, yet the effective design of transformer architectures for MDM…
MMTB: Evaluating Terminal Agents on Multimedia-File Tasks
Chiyeong Heo, Jaechang Kim, Junhyuk Kwon +4
Terminals provide a powerful interface for AI agents by exposing diverse tools for automating complex workflows, yet existing terminal-agent benchmarks largely focus on tasks groun…
KVoiceBench, KOpenAudioBench, and KMMAU: Agent-Driven Korean Speech Benchmarks for Evaluating SpeechLMs
Haechan Kim, Seungjun Chung, Inkyu Park +2
Speech language models (SpeechLMs) have achieved substantial progress by extending large language models (LLMs) to the speech modality. However, SpeechLM evaluation remains heavily…
Raon-Speech Technical Report
Beomsoo Kim, Changho Choi, Dohyun Kim +23
We present Raon-Speech, a top-performing 9B-parameter speech language model (SpeechLM) for English and Korean speech understanding, answering, and generation, and Raon-SpeechChat,…
Exploring Multimodal Perception in Large Language Models Through Perceptual Strength Ratings
Jonghyun Lee, Dojun Park, Jiwoo Lee +2
This study investigated whether multimodal large language models can achieve human-like sensory grounding by examining their ability to capture perceptual strength ratings across s…