activity
20192025
collaborators
Showing cs.CLShow all

11 papers · 1 filter

cs.CL2026

Negative Self-Distillation: Learning to Reason by Avoiding Flaws

Rongcan Pei, Zhepei Wei, Shuyao Xu +3

On-Policy Self-Distillation (OPSD) has emerged as a popular paradigm for large language model (LLM) self-improvement, allowing models to act as their own teachers by leveraging pri…

cs.CL2025

Separate the Wheat from the Chaff: Winnowing Down Divergent Views in Retrieval Augmented Generation

Song Wang, Zihan Chen, Peng Wang +5

Retrieval-augmented generation (RAG) enhances large language models (LLMs) by integrating external knowledge sources to address their limitations in accessing up-to-date or special…

cs.CL2025

Aligning Large Language Models via Fully Self-Synthetic Data

Shangjian Yin, Zhepei Wei, Xinyu Zhu +2

Traditional reinforcement learning from human feedback (RLHF) for large language models (LLMs) relies on expensive human-annotated datasets, while Reinforcement Learning from AI Fe…

cs.CL2025

TruthRL: Incentivizing Truthful LLMs via Reinforcement Learning

Zhepei Wei, Xiao Yang, Kai Sun +12

While large language models (LLMs) have demonstrated strong performance on factoid question answering, they are still prone to hallucination and untruthful responses, particularly…

cs.CL2025

AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism

Zhepei Wei, Wei-Lin Chen, Xinyu Zhu +1

Large language models (LLMs) are increasingly used for long-content generation (e.g., long Chain-of-Thought reasoning) where decoding efficiency becomes a critical bottleneck: Auto…

cs.CL2025

The Surprising Effectiveness of Negative Reinforcement in LLM Reasoning

Xinyu Zhu, Mengzhou Xia, Zhepei Wei +3

Reinforcement learning with verifiable rewards (RLVR) is a promising approach for training language models (LMs) on reasoning tasks that elicit emergent long chains of thought (CoT…