9 papers · 1 filter
DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects
Yi Shu, Tianyu Peng, Yingzhuo Deng +5
Current end-to-end speech dialogue models are primarily optimized for mainstream languages and remain limited in low-resource dialect scenarios due to the scarcity of dialect speec…
TokAlign++: Advancing Vocabulary Adaptation via Better Token Alignment
Chong Li, Yingzhuo Deng, Wen Yang +2
Tokenization is a foundational step in the text process of Large Language Models (LLMs). Texts must be first tokenized into token IDs, which are then input to LLMs. Inefficient tok…
OpenS2S: Advancing Fully Open-Source End-to-End Empathetic Large Speech Language Model
Chen Wang, Tianyu Peng, Wen Yang +8
Empathetic interaction is a cornerstone of human-machine communication, due to the need for understanding speech enriched with paralinguistic cues and generating emotional and expr…
Parallel Scaling Law: Unveiling Reasoning Generalization through A Cross-Linguistic Perspective
Wen Yang, Junhong Wu, Chong Li +2
Recent advancements in Reinforcement Post-Training (RPT) have significantly enhanced the capabilities of Large Reasoning Models (LRMs), sparking increased interest in the generaliz…
Implicit Cross-Lingual Rewarding for Efficient Multilingual Preference Alignment
Wen Yang, Junhong Wu, Chen Wang +2
Direct Preference Optimization (DPO) has become a prominent method for aligning Large Language Models (LLMs) with human preferences. While DPO has enabled significant progress in a…
Language Imbalance Driven Rewarding for Multilingual Self-improving
Wen Yang, Junhong Wu, Chen Wang +2
Large Language Models (LLMs) have achieved state-of-the-art performance across numerous tasks. However, these advancements have predominantly benefited "first-class" languages such…