2 papers
cs.SD2026
Joint Analysis of Latent Dimensionality and Frame Rate in Continuous Audio Encoders
Kyudan Jung, Sehyun Lee, Song-ha Jo +4
Continuous audio encoders compress audio along feature and time axes through latent width and frame rate, but their joint effect on downstream performance remains unclear. We train…
cs.SD2026
Where Does the Sound Go? Tracing Acoustic Information Loss in Audio-Conditioned LLMs
Song-ha Jo, Sehyun Lee, Soyoon Kim +2
Audio-conditioned language models often underuse acoustic cues such as prosody, emotion, and non-speech sounds, raising the question of whether ASR-supervised frontends discard thi…