1 paper
KiHyun Nam, Jungwoo Heo, Siu Bae +2
As audio-first agents become increasingly common in physical AI, conversational robots, and screenless wearables, audio large language models (audio-LLMs) must integrate speaker-sp…