works on

From the 1 of 14 linked papers with an AI index.

activity
20242026
collaborators

14 papers

cs.RO2026

Kepler-Encoder-v0.1: Towards a Multimodal Embedding Model for Robots

Ishneet Sukhvinder Singh, Dhanoosh Pooranakumaran, Alex Nguyen +1

The paper introduces Kepler-Encoder-v0.1, a self‑supervised multimodal encoder that fuses vision, proprioception, and force/torque data into a shared latent space, enabling a visio…

eess.AS2026

Emotional Dimension Control in Language Model-Based Text-to-Speech: Spanning a Broad Spectrum of Human Emotions

Kun Zhou, You Zhang, Dianwen Ng +3

Emotional text-to-speech (TTS) systems sturggle to capture the full spectrum of human emotions due to the inherent complexity of emotional expressions and the limited coverage of e…

cs.SD2025

Bona fide Cross Testing Reveals Weak Spot in Audio Deepfake Detection Systems

Chin Yuen Kwok, Jia Qi Yip, Zhen Qiu +2

Audio deepfake detection (ADD) models are commonly evaluated using datasets that combine multiple synthesizers, with performance reported as a single Equal Error Rate (EER). Howeve…

cs.CL2025

Improving Synthetic Data Training for Contextual Biasing Models with a Keyword-Aware Cost Function

Chin Yuen Kwok, Jia Qi Yip, Eng Siong Chng

Rare word recognition can be improved by adapting ASR models to synthetic data that includes these words. Further improvements can be achieved through contextual biasing, which tra…

cs.CL2025

Efficient Trie-based Biasing using K-step Prediction for Rare Word Recognition

Chin Yuen Kwok, Jia Qi yip

Contextual biasing improves rare word recognition of ASR models by prioritizing the output of rare words during decoding. A common approach is Trie-based biasing, which gives "bonu…

cs.SD2025

Analysis of Speaker Verification Performance Trade-offs with Neural Audio Codec Transmission

Nirmalya Mallick Thakur, Jia Qi Yip, Eng Siong Chng

Neural audio codecs (NACs) have made significant advancements in recent years and are rapidly being adopted in many audio processing pipelines. However, they can introduce audio di…