works on

From the 1 of 17 linked papers with an AI index.

activity
20242026
collaborators
Showing 2025Show all

10 papers · 1 filter

cs.SD2025

Bona fide Cross Testing Reveals Weak Spot in Audio Deepfake Detection Systems

Chin Yuen Kwok, Jia Qi Yip, Zhen Qiu +2

Audio deepfake detection (ADD) models are commonly evaluated using datasets that combine multiple synthesizers, with performance reported as a single Equal Error Rate (EER). Howeve…

cs.CL2025

Improving Synthetic Data Training for Contextual Biasing Models with a Keyword-Aware Cost Function

Chin Yuen Kwok, Jia Qi Yip, Eng Siong Chng

Rare word recognition can be improved by adapting ASR models to synthetic data that includes these words. Further improvements can be achieved through contextual biasing, which tra…

cs.CL2025

Efficient Trie-based Biasing using K-step Prediction for Rare Word Recognition

Chin Yuen Kwok, Jia Qi yip

Contextual biasing improves rare word recognition of ASR models by prioritizing the output of rare words during decoding. A common approach is Trie-based biasing, which gives "bonu…

cs.SD2025

Analysis of Speaker Verification Performance Trade-offs with Neural Audio Codec Transmission

Nirmalya Mallick Thakur, Jia Qi Yip, Eng Siong Chng

Neural audio codecs (NACs) have made significant advancements in recent years and are rapidly being adopted in many audio processing pipelines. However, they can introduce audio di…

cs.CL2025

Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks

Chien-yu Huang, Wei-Chih Chen, Shu-wen Yang +77

Multimodal foundation models, such as Gemini and ChatGPT, have revolutionized human-machine interactions by seamlessly integrating various forms of data. Developing a universal spo…

eess.AS2025

Speechless: Speech Instruction Training Without Speech for Low Resource Languages

Alan Dao, Dinh Bach Vu, Huy Hoang Ha +6

The rapid growth of voice assistants powered by large language models (LLM) has highlighted a need for speech instruction data to train these systems. Despite the abundance of spee…