activity
20242026
collaborators

8 papers

cs.LG2026

Cubit: Token Mixer with Kernel Ridge Regression

Chuanyang Zheng, Jiankai Sun, Yihang Gao +6

Since its introduction in 2017, the Transformer has become one of the most widely adopted architectures in modern deep learning. Despite extensive efforts to improve positional enc…

cs.LG2026

GeoNorm: Unify Pre-Norm and Post-Norm with Geodesic Optimization

Chuanyang Zheng, Jiankai Sun, Yihang Gao +11

The placement of normalization layers, specifically Pre-Norm and Post-Norm, remains an open question in Transformer architecture design. In this work, we rethink these approaches t…

cs.CL2025

SAS: Simulated Attention Score

Chuanyang Zheng, Jiankai Sun, Yihang Gao +12

The attention mechanism is a core component of the Transformer architecture. Various methods have been developed to compute attention scores, including multi-head attention (MHA),…

cs.CL2025

Understanding the Mixture-of-Experts with Nadaraya-Watson Kernel

Chuanyang Zheng, Jiankai Sun, Yihang Gao +13

Mixture-of-Experts (MoE) has become a cornerstone in recent state-of-the-art large language models (LLMs). Traditionally, MoE relies on as the router score funct…

cs.CL2025

SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator

Guoxuan Chen, Han Shi, Jiawei Li +7

Large Language Models (LLMs) have exhibited exceptional performance across a spectrum of natural language processing tasks. However, their substantial sizes pose considerable chall…

cs.CL2025

Self-Adjust Softmax

Chuanyang Zheng, Yihang Gao, Guoxuan Chen +7

The softmax function is crucial in Transformer attention, which normalizes each row of the attention scores with summation to one, achieving superior performances over other altern…