3 papers
cs.CL2024
MLKD-BERT: Multi-level Knowledge Distillation for Pre-trained Language Models
Ying Zhang, Ziheng Yang, Shufan Ji
Knowledge distillation is an effective technique for pre-trained language model compression. Although existing knowledge distillation methods perform well for the most typical mode…
cs.CL2023
Bidirectional Transformer Reranker for Grammatical Error Correction
Ying Zhang, Hidetaka Kamigaito, Manabu Okumura
Pre-trained seq2seq models have achieved state-of-the-art results in the grammatical error correction task. However, these models still suffer from a prediction bias due to their u…
cs.SD2023
Improving Prosody for Cross-Speaker Style Transfer by Semi-Supervised Style Extractor and Hierarchical Modeling in Speech Synthesis
Chunyu Qiang, Peng Yang, Hao Che +3
Cross-speaker style transfer in speech synthesis aims at transferring a style from source speaker to synthesized speech of a target speaker's timbre. In most previous methods, the…