collaborators
Showing cs.CLShow all

8 papers · 1 filter

cs.CL2025

Why Not Transform Chat Large Language Models to Non-English?

Xiang Geng, Ming Zhu, Jiahuan Li +14

The scarcity of non-English data limits the development of non-English large language models (LLMs). Transforming English-centric LLMs to non-English has been identified as an effe…

cs.CL2025

Investigating Numerical Translation with Large Language Models

Wei Tang, Jiawei Yu, Yuang Li +5

The inaccurate translation of numbers can lead to significant security issues, ranging from financial setbacks to medical inaccuracies. While large language models (LLMs) have made…

cs.CL2024

"I've Heard of You!": Generate Spoken Named Entity Recognition Data for Unseen Entities

Jiawei Yu, Xiang Geng, Yuang Li +8

Spoken named entity recognition (NER) aims to identify named entities from speech, playing an important role in speech processing. New named entities appear every day, however, ann…

cs.CL2024

Hard-Synth: Synthesizing Diverse Hard Samples for ASR using Zero-Shot TTS and LLM

Jiawei Yu, Yuang Li, Xiaosong Qiao +6

Text-to-speech (TTS) models have been widely adopted to enhance automatic speech recognition (ASR) systems using text-only corpora, thereby reducing the cost of labeling real speec…

cs.CL2024

From Handcrafted Features to LLMs: A Brief Survey for Machine Translation Quality Estimation

Haofei Zhao, Yilun Liu, Shimin Tao +6

Machine Translation Quality Estimation (MTQE) is the task of estimating the quality of machine-translated text in real time without the need for reference translations, which is of…

cs.CL2024

Large Language Model Should Understand Pinyin for Chinese ASR Error Correction

Yuang Li, Xiaosong Qiao, Xiaofeng Zhao +4

Large language models can enhance automatic speech recognition systems through generative error correction. In this paper, we propose Pinyin-enhanced GEC, which leverages Pinyi, th…