6 papers
WuNeng: Hybrid State with Attention
Liu Xiao, Li Zhiyuan, Lin Yueyu
The WuNeng architecture introduces a novel approach to enhancing the expressivity and power of large language models by integrating recurrent neural network (RNN)-based RWKV-7 with…
Cross-attention for State-based model RWKV-7
Liu Xiao, Li Zhiyuan, Lin Yueyu
We introduce CrossWKV, a novel cross-attention mechanism for the state-based RWKV-7 model, designed to enhance the expressive power of text-to-image generation. Leveraging RWKV-7's…
Millions of States: Designing a Scalable MoE Architecture with RWKV-7 Meta-learner
Liu Xiao, Li Zhiyuan, Lin Yueyu
State-based sequence models like RWKV-7 offer a compelling alternative to Transformer architectures, achieving linear complexity while demonstrating greater expressive power in sho…
State Tuning: State-based Test-Time Scaling on RWKV-7
Liu Xiao, Li Zhiyuan, Lin Yueyu
Test-time scaling has emerged as a prominent research direction in machine learning, enabling models to enhance their expressive capabilities during inference.Transformers, renowne…
RWKVTTS: Yet another TTS based on RWKV-7
Lin yueyu, Liu Xiao
Human-AI interaction thrives on intuitive and efficient interfaces, among which voice stands out as a particularly natural and accessible modality. Recent advancements in transform…
ARWKV: Pretrain is not what we need, an RNN-Attention-Based Language Model Born from Transformer
Lin Yueyu, Li Zhiyuan, Peter Yue +1
As is known, hybrid quadratic and subquadratic attention models in multi-head architectures have surpassed both Transformer and Linear RNN models , with these works primarily focus…