3 papers
cs.SD2026
WASIL: In-the-Wild Arabic Spoken Interactions with LLMs
Zien Sheikh Ali, Hamdy Mubarak, Soon-Gyo Jung +3
Large Language Models (LLMs) voice assistants are commonly built as cascaded Automatic Speech recognition (ASR) to LLM systems, where recognition errors can distort user intent. Di…
cs.CL2026
OASIS: A Multilingual and Multimodal Dataset for Culturally Grounded Spoken Visual QA
Firoj Alam, Ali Ezzat Shahroor, Md. Arid Hasan +8
Large-scale multimodal models achieve strong results on tasks like Visual Question Answering (VQA), but they are often limited when queries require cultural and visual information,…
cs.SD2026
MENASpeechBank: A Reference Voice Bank with Persona-Conditioned Multi-Turn Conversations for AudioLLMs
Zien Sheikh Ali, Hunzalah Hassan Bhatti, Rabindra Nath Nandi +2
Audio large language models (AudioLLMs) enable instruction-following over speech and general audio, but progress is increasingly limited by the lack of diverse, conversational, ins…