collaborators

10 papers

cs.AI2025

LADY: Linear Attention for Autonomous Driving Efficiency without Transformers

Jihao Huang, Xi Xia, Zhiyuan Li +4

End-to-end autonomous driving has emerged as a promising paradigm. However, state-of-the-art methods rely heavily on Transformer architectures. The inherent quadratic complexity of…

cs.CV2025

PersonaLive! Expressive Portrait Image Animation for Live Streaming

Zhiyuan Li, Chi-Man Pun, Chen Fang +2

Current diffusion-based portrait animation models predominantly focus on enhancing visual quality and expression realism, while overlooking generation latency and real-time perform…

cs.LG2025

Provable Benefits of Sinusoidal Activation for Modular Addition

Tianlong Huang, Zhiyuan Li

This paper studies the role of activation functions in learning modular addition with two-layer neural networks. We first establish a sharp expressivity gap: sine MLPs admit width-…

cs.CL2025

Kimi Linear: An Expressive, Efficient Attention Architecture

Kimi Team, Yu Zhang, Zongyu Lin +57

We introduce Kimi Linear, a hybrid linear attention architecture that, for the first time, outperforms full attention under fair comparisons across various scenarios -- including s…

cs.CL2025

WuNeng: Hybrid State with Attention

Liu Xiao, Li Zhiyuan, Lin Yueyu

The WuNeng architecture introduces a novel approach to enhancing the expressivity and power of large language models by integrating recurrent neural network (RNN)-based RWKV-7 with…

cs.CV2025

Cross-attention for State-based model RWKV-7

Liu Xiao, Li Zhiyuan, Lin Yueyu

We introduce CrossWKV, a novel cross-attention mechanism for the state-based RWKV-7 model, designed to enhance the expressive power of text-to-image generation. Leveraging RWKV-7's…