11 papers
Full-Stack FP4: Stable LLM Pretraining with Quantized Projections, Optimizers, and Attention
Siyu Ding, Mingchuan Ma, Jiabo Tong +3
Recent NVFP4 pretraining work has primarily optimized Transformer linear projections, leaving persistent optimizer states, optimizer computation, and low-precision attention forwar…
NeuroCogMap Reveals Cognitive Organization of Large Language Models
Zhongxiang Sun, Haolang Lu, Qiang Ma +11
Understanding how complex cognitive functions are organized within artificial systems is central to interpreting large language models (LLMs) and relating them to biological cognit…
Efficient Sparse Selective-Update RNNs for Long-Range Sequence Modeling
Bojian Yin, Shurong Wang, Haoyu Tan +3
Real-world sequential signals, such as audio or video, contain critical information that is often embedded within long periods of silence or noise. While recurrent neural networks…
Autoregressive Image Generation with Randomized Parallel Decoding
Haopeng Li, Jinyue Yang, Guoqi Li +1
We introduce ARPG, a novel visual Autoregressive model that enables Randomized Parallel Generation, addressing the inherent limitations of conventional raster-order approaches, whi…
PretrainZero: Reinforcement Active Pretraining
Xingrun Xing, Zhiyuan Fan, Jie Lou +3
Mimicking human behavior to actively learning from general experience and achieve artificial general intelligence has always been a human dream. Recent reinforcement learning (RL)…
Breaking the Modality Wall: Time-step Mixup for Efficient Spiking Knowledge Transfer from Static to Event Domain
Yuqi Xie, Shuhan Ye, Yi Yu +7
The integration of event cameras and spiking neural networks (SNNs) promises energy-efficient visual intelligence, yet scarce event data and the sparsity of DVS outputs hinder effe…