most citedARWKV: Pretrain is not what we need, an RNN-Attention-Based Language Model Born from Transformer

1 citations · 1 across the 6 of their papers we have counts for

collaborators

7 papers

cs.CL2026

DREAMSTATE: Diffusing States and Parameters for Recurrent Large Language Models

Liu Xiao

Modern Recurrent Neural Networks (RNNs), such as RWKV, are distinguished by their powerful short-range modeling capabilities and efficient fixed-size states, which constitute a cor…

cs.CL2025

WuNeng: Hybrid State with Attention

Liu Xiao, Li Zhiyuan, Lin Yueyu

The WuNeng architecture introduces a novel approach to enhancing the expressivity and power of large language models by integrating recurrent neural network (RNN)-based RWKV-7 with…

cs.CV2025

Cross-attention for State-based model RWKV-7

Liu Xiao, Li Zhiyuan, Lin Yueyu

We introduce CrossWKV, a novel cross-attention mechanism for the state-based RWKV-7 model, designed to enhance the expressive power of text-to-image generation. Leveraging RWKV-7's…

cs.LG2025

Millions of States: Designing a Scalable MoE Architecture with RWKV-7 Meta-learner

Liu Xiao, Li Zhiyuan, Lin Yueyu

State-based sequence models like RWKV-7 offer a compelling alternative to Transformer architectures, achieving linear complexity while demonstrating greater expressive power in sho…

cs.SD2025

RWKVTTS: Yet another TTS based on RWKV-7

Lin yueyu, Liu Xiao

Human-AI interaction thrives on intuitive and efficient interfaces, among which voice stands out as a particularly natural and accessible modality. Recent advancements in transform…

cs.LG2025

BlackGoose Rimer: Harnessing RWKV-7 as a Simple yet Superior Replacement for Transformers in Large-Scale Time Series Modeling

Li weile, Liu Xiao

Time series models face significant challenges in scaling to handle large and complex datasets, akin to the scaling achieved by large language models (LLMs). The unique characteris…