60 citations · 70 across the 7 of their papers we have counts for
Showing cs.SDShow all
2 papers · 1 filter
cs.SD2026
AudioSpan: Spanning the Duration and Depth of Audio Comprehension
Wen Huang, Yunfei Chu, Meng Gao +2
General audio comprehension now covers speech, sound, and music over durations from seconds to hours, driven by large audio-language models (LALMs) that are increasingly omni-modal…
cs.SD2023
LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT
Zhihao Du, Jiaming Wang, Qian Chen +12
Generative Pre-trained Transformer (GPT) models have achieved remarkable performance on various natural language processing tasks, and have shown great potential as backbones for a…