2 papers
cs.SD2026
Music I Care About: Automated Multimodal Benchmarking of LLM Music Perception Skills on (Almost) Any Music
Tomáš Sourada, Katia Vendrame, Jan HajiÄ
Music represents a cornerstone of human culture, existing digitally across diverse modalities, including audio, symbolic encodings (e.g., MIDI, MusicXML), and sheet music. Despite…
cs.CL2026
Joint Speech and Text Training for LLM-Based End-to-End Spoken Dialogue State Tracking
Katia Vendrame, Bolaji Yusuf, Santosh Kesiraju +3
End-to-end spoken dialogue state tracking (DST) is made difficult by the tandem of having to handle speech input and data scarcity. Combining speech foundation encoders and large l…