1 paper · 1 filter
Tony Lee, Haoqin Tu, Chi Heem Wong +6
Evaluations of audio-language models (ALMs) -- multimodal models that take interleaved audio and text as input and output text -- are hindered by the lack of standardized benchmark…