collaborators

6 papers

cs.CV2026

The Lipreading Gap: Do VSR Models Perceive Visual Speech Like Human Lipreaders?

Rishabh Jain, Naomi Harte

Visual speech recognition (VSR) models now surpass human lipreaders on benchmarks, but do such gains establish human-like visual speech perception? To explore this, we compare thre…

eess.AS2026

VisG AV-HuBERT: Viseme-Guided AV-HuBERT

Aristeidis Papadopoulos, Rishabh Jain, Naomi Harte

Audio-Visual Speech Recognition (AVSR) systems nowadays integrate Large Language Model (LLM) decoders with transformer-based encoders, achieving state-of-the-art results. However,…

cs.AI2023

Data Center Audio/Video Intelligence on Device (DAVID) -- An Edge-AI Platform for Smart-Toys

Gabriel Cosache, Francisco Salgado, Cosmin Rotariu +3

An overview is given of the DAVID Smart-Toy platform, one of the first Edge AI platform designs to incorporate advanced low-power data processing by neural inference models co-loca…

cs.HC2023

Synthetic Speaking Children -- Why We Need Them and How to Make Them

Muhammad Ali Farooq, Dan Bigioi, Rishabh Jain +3

Contemporary Human Computer Interaction (HCI) research relies primarily on neural network models for machine vision and speech understanding of a system user. Such models require e…

cs.CL2023

A comparative analysis between Conformer-Transducer, Whisper, and wav2vec2 for improving the child speech recognition

Andrei Barcovschi, Rishabh Jain, Peter Corcoran

Automatic Speech Recognition (ASR) systems have progressed significantly in their performance on adult speech data; however, transcribing child speech remains challenging due to th…

cs.SD2023

Improved Child Text-to-Speech Synthesis through Fastpitch-based Transfer Learning

Rishabh Jain, Peter Corcoran

Speech synthesis technology has witnessed significant advancements in recent years, enabling the creation of natural and expressive synthetic speech. One area of particular interes…