activity
20242026
collaborators
Showing cs.CLShow all

11 papers · 1 filter

cs.CL2026

EDGE: Experience-Distillation for Guided Exploration in Agentic Reinforcement Learning

Can Xie, Yuyi Zhou, Wen Yang +5

Reinforcement learning with outcome-based objectives such as GRPO enables LLM-based agents to solve complex, long-horizon tasks, yet the reusable exploration patterns embedded in i…

cs.CL2026

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects

Yi Shu, Tianyu Peng, Yingzhuo Deng +5

Current end-to-end speech dialogue models are primarily optimized for mainstream languages and remain limited in low-resource dialect scenarios due to the scarcity of dialect speec…

cs.CL2026

TokAlign++: Advancing Vocabulary Adaptation via Better Token Alignment

Chong Li, Yingzhuo Deng, Wen Yang +2

Tokenization is a foundational step in the text process of Large Language Models (LLMs). Texts must be first tokenized into token IDs, which are then input to LLMs. Inefficient tok…

cs.CL2025

Parallel Scaling Law: Unveiling Reasoning Generalization through A Cross-Linguistic Perspective

Wen Yang, Junhong Wu, Chong Li +2

Recent advancements in Reinforcement Post-Training (RPT) have significantly enhanced the capabilities of Large Reasoning Models (LRMs), sparking increased interest in the generaliz…

cs.CL2025

OpenS2S: Advancing Fully Open-Source End-to-End Empathetic Large Speech Language Model

Chen Wang, Tianyu Peng, Wen Yang +8

Empathetic interaction is a cornerstone of human-machine communication, due to the need for understanding speech enriched with paralinguistic cues and generating emotional and expr…

cs.CL2025

Implicit Cross-Lingual Rewarding for Efficient Multilingual Preference Alignment

Wen Yang, Junhong Wu, Chen Wang +2

Direct Preference Optimization (DPO) has become a prominent method for aligning Large Language Models (LLMs) with human preferences. While DPO has enabled significant progress in a…