3 papers
cs.AI2026
Full-bandwidth transformer
Xi Wang, Ziyang Cai, Zheng Zhan +5
Autoregressive transformers compute along two axes: horizontally across generated tokens, and vertically through model depth. Dense attention gives each token broad horizontal acce…
cs.CL2026
Hierarchical Latent Prediction for Language Models
Chang Shi, Tim Pearce, Manan Tomar +2
While standard Next-Token Prediction (NTP) lays the foundation of language model pre- training, its teacher-forced training paradigm may not be optimal for long-horizon reasoning a…
cs.LG2024
Scaling Laws for Pre-training Agents and World Models
Tim Pearce, Tabish Rashid, Dave Bignell +3
The performance of embodied agents has been shown to improve by increasing model parameters, dataset size, and compute. This has been demonstrated in domains from robotics to video…