4 papers
CallShield: Secure Caller Authentication over Real-Time Audio Channels
Mouna Rabh, Yazan Boshmaf, Mashael Alsabah +3
We present CallShield, the first caller identity authentication system that operates entirely at the audio layer, without relying on speech transcription, internet connectivity, or…
From Words to Waves: Analyzing Concept Formation in Speech and Text-Based Foundation Models
Asım Ersoy, Basel Mousi, Shammur Chowdhury +3
The emergence of large language models (LLMs) has demonstrated that systems trained solely on text can acquire extensive world knowledge, develop reasoning capabilities, and intern…
SpokenNativQA: Multilingual Everyday Spoken Queries for LLMs
Firoj Alam, Md Arid Hasan, Shammur Absar Chowdhury
Large Language Models (LLMs) have demonstrated remarkable performance across various disciplines and tasks. However, benchmarking their capabilities with multilingual spoken querie…
BnTTS: Few-Shot Speaker Adaptation in Low-Resource Setting
Mohammad Jahid Ibna Basher, Md Kowsher, Md Saiful Islam +8
This paper introduces BnTTS (Bangla Text-To-Speech), the first framework for Bangla speaker adaptation-based TTS, designed to bridge the gap in Bangla speech synthesis using minima…