collaborators

5 papers

cs.CL2026

EDGE: Experience-Distillation for Guided Exploration in Agentic Reinforcement Learning

Can Xie, Yuyi Zhou, Wen Yang +5

Reinforcement learning with outcome-based objectives such as GRPO enables LLM-based agents to solve complex, long-horizon tasks, yet the reusable exploration patterns embedded in i…

cs.CL2026

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects

Yi Shu, Tianyu Peng, Yingzhuo Deng +5

Current end-to-end speech dialogue models are primarily optimized for mainstream languages and remain limited in low-resource dialect scenarios due to the scarcity of dialect speec…

cs.AI2026

AMR-SD: Asymmetric Meta-Reflective Self-Distillation for Token-Level Credit Assignment

Zhenlin Wei, Pu Jian, Yingzhuo Deng +6

The alignment of Large Language Models (LLMs) for complex reasoning heavily relies on Reinforcement Learning with Verifiable Rewards (RLVR). However, standard algorithms like GRPO…

cs.CL2026

TokAlign++: Advancing Vocabulary Adaptation via Better Token Alignment

Chong Li, Yingzhuo Deng, Wen Yang +2

Tokenization is a foundational step in the text process of Large Language Models (LLMs). Texts must be first tokenized into token IDs, which are then input to LLMs. Inefficient tok…

cs.CL2025

Group then Scale: Dynamic Mixture-of-Experts Multilingual Language Model

Chong Li, Yingzhuo Deng, Jiajun Zhang +1

The curse of multilinguality phenomenon is a fundamental problem of multilingual Large Language Models (LLMs), where the competition between massive languages results in inferior p…