1 paper · 1 filter
Apoorva Kulkarni, Kaousheik Jayakumar, Sreyan Ghosh +3
Large audio-language models (LALMs) describe audio at the clip level but cannot assign timestamps to the events, speakers, or sounds they identify. Despite being essential for down…