activity
20202025
collaborators

5 papers

eess.AS2025

Target Speaker Lipreading by Audio-Visual Self-Distillation Pretraining and Speaker Adaptation

Jing-Xuan Zhang, Tingzhi Mao, Longjiang Guo +2

Lipreading is an important technique for facilitating human-computer interaction in noisy environments. Our previously developed self-supervised learning method, AV2vec, which leve…

cs.CL2024

Lightweight Transducer Based on Frame-Level Criterion

Genshun Wan, Mengzhi Wang, Tingzhi Mao +2

The transducer model trained based on sequence-level criterion requires a lot of memory due to the generation of the large probability matrix. We proposed a lightweight transducer…

cs.SD2020

Enriching Under-Represented Named-Entities To Improve Speech Recognition Performance

Tingzhi Mao, Yerbolat Khassanov, Van Tung Pham +4

Automatic speech recognition (ASR) for under-represented named-entity (UR-NE) is challenging due to such named-entities (NE) have insufficient instances and poor contextual coverag…

eess.AS2020

The NTU-AISG Text-to-speech System for Blizzard Challenge 2020

Haobo Zhang, Tingzhi Mao, Haihua Xu +1

We report our NTU-AISG Text-to-speech (TTS) entry systems for the Blizzard Challenge 2020 in this paper. There are two TTS tasks in this year's challenge, one is a Mandarin TTS tas…

eess.AS2020

Approaches to Improving Recognition of Underrepresented Named Entities in Hybrid ASR Systems

Tingzhi Mao, Yerbolat Khassanov, Van Tung Pham +3

In this paper, we present a series of complementary approaches to improve the recognition of underrepresented named entities (NE) in hybrid ASR systems without compromising overall…