Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
When Audio-LLMs Don't Listen: A Cross-Linguistic Study of Modality Arbitration
Jayadev Billa
When audio and text conflict, speech-enabled language models follow text far more often than they do when arbitrating between two conflicting text sources, even under explicit inst…
cs.CL2026
Modality Collapse as Mismatched Decoding: Information-Theoretic Limits of Multimodal LLMs
Jayadev Billa
Numerous studies have shown that multimodal LLMs process speech and images well but fail in non-intuitive ways rendering trivial tasks such as object counting unreliable. We invest…
cs.CL2026
The Cascade Equivalence Hypothesis: When Do Speech LLMs Behave Like ASRLLM Pipelines?
Jayadev Billa
Speech LLMs are widely understood to be better than ASRLLM cascades since they have access to the audio directly, and not just the transcript. In this paper, we presen…