60 citations · 127 across the 20 of their papers we have counts for
4 papers · 1 filter
OmniEcho: Audio-Visual Spatial Understanding for Omni-Modal Embodied Agents
Ruixun Liu, Yuxuan Wang, Jiacheng Xie +10
Humans can effortlessly localize the direction of a sound source and integrate it with visual cues for reasoning, yet this remains challenging for embodied agents. In particular, i…
LongAudioSpan: Spanning the Duration and Depth of Audio Comprehension
Wen Huang, Yunfei Chu, Meng Gao +2
General audio comprehension now covers speech, sound, and music over durations from seconds to hours, driven by large audio-language models (LALMs) that are increasingly omni-modal…
Qwen-Music Technical Report
Jin Xu, Kangdi Wang, Ruibin Yuan +24
We introduce Qwen-Music, a music generation model that produces high-fidelity songs with complete vocals. It supports text-to-music generation from descriptions, lyrics, and musica…
LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT
Zhihao Du, Jiaming Wang, Qian Chen +12
Generative Pre-trained Transformer (GPT) models have achieved remarkable performance on various natural language processing tasks, and have shown great potential as backbones for a…