activity
20242026
collaborators
Showing cs.CLShow all

13 papers · 1 filter

cs.CL2026

Coupling Experts and Routers in Mixture-of-Experts via an Auxiliary Loss

Ang Lv, Jin Ma, Yiyuan Ma +1

Mixture-of-Experts (MoE) models lack explicit constraints to ensure the router's decisions align well with the experts' capabilities, which ultimately limits model performance. To…

cs.CL2025

Autonomy-of-Experts Models

Ang Lv, Ruobing Xie, Yining Qian +5

Mixture-of-Experts (MoE) models mostly use a router to assign tokens to specific expert modules, activating only partial parameters and often outperforming dense models. We argue t…

cs.CL2025

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason

Ang Lv, Ruobing Xie, Xingwu Sun +2

Recent studies on post-training large language models (LLMs) for reasoning through reinforcement learning (RL) typically focus on tasks that can be accurately verified and rewarded…

cs.CL2025

Language Models "Grok" to Copy

Ang Lv, Ruobing Xie, Xingwu Sun +2

We examine the pre-training dynamics of language models, focusing on their ability to copy text from preceding context--a fundamental skill for various LLM applications, including…

cs.CL2025

More Expressive Attention with Negative Weights

Ang Lv, Ruobing Xie, Shuaipeng Li +5

We propose a novel attention mechanism, named Cog Attention, that enables attention weights to be negative for enhanced expressiveness, which stems from two key factors: (1) Cog At…

cs.CL2024

HoPE: A Novel Positional Encoding Without Long-Term Decay for Enhanced Context Awareness and Extrapolation

Yuhan Chen, Ang Lv, Jian Luan +2

Many positional encodings (PEs) are designed to exhibit long-term decay, based on an entrenched and long-standing inductive opinion: tokens farther away from the current position c…