9 papers
Understanding the Effects of Distractors on Reasoning Vision-Language Models
Jiyun Bae, Hyunjong Ok, Sangwoo Mo +1
How does irrelevant information (i.e., distractors) affect test-time scaling in vision-language models (VLMs)? Prior work on text-only language models has shown that textual distra…
Lost in the Prompt Order: Revealing the Limitations of Causal Attention in Language Models
Hyunjong Ok, Jaeho Lee
Large language models exhibit surprising sensitivity to the structure of the prompt, but the mechanisms underlying this sensitivity remain poorly understood. In this work, we condu…
Speculative End-Turn Detector for Efficient Speech Chatbot Assistant
Hyunjong Ok, Suho Yoo, Jaeho Lee
Spoken dialogue systems powered by large language models have demonstrated remarkable abilities in understanding human speech and generating appropriate spoken responses. However,…
TempCore: Are Video QA Benchmarks Temporally Grounded? A Frame Selection Sensitivity Analysis and Benchmark
Hyunjong Ok, Jaeho Lee
Vision-language models (VLMs) can ingest only a limited number of video frames, making frame selection a practical necessity. But do current Video QA benchmarks genuinely require t…
AuditoryBench++: Can Language Models Understand Auditory Knowledge without Hearing?
Hyunjong Ok, Suho Yoo, Hyeonjun Kim +1
Even without directly hearing sounds, humans can effortlessly reason about auditory properties, such as pitch, loudness, or sound-source associations, drawing on auditory commonsen…
S2Cap: A Benchmark and a Baseline for Singing Style Captioning
Hyunjong Ok, Jaeho Lee
Singing voices contain much richer information than common voices, including varied vocal and acoustic properties. However, current open-source audio-text datasets for singing voic…