3 citations · 8 across the 9 of their papers we have counts for
1 paper · 1 filter
Kartik Hegde, Rehana Mahfuz, Yinyi Guo +1
Current audio captioning relies on supervised learning with paired audio-caption data, which is costly to curate and may not reflect human preferences in real-world scenarios. To a…