Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Dynamic Short Convolutions Improve Transformers
Oliver Sieberling, Bharat Runwal, Rameswar Panda +1
Transformers have become the dominant architecture for large language models, largely due to the scalability and flexibility of attention, feed-forward layers, residual connections…
cs.LG2026
PRISM: Demystifying Retention and Interaction in Mid-Training
Bharat Runwal, Ashish Agrawal, Anurag Roy +1
We present PRISM, a comprehensive empirical study of mid-training design choices for large language models. Through controlled experiments across seven base models spanning four fa…
cs.LG2025
FlashFormer: Whole-Model Kernels for Efficient Low-Batch Inference
Aniruddha Nrusimha, William Brandon, Mayank Mishra +4
The size and compute characteristics of modern large language models have led to an increased interest in developing specialized kernels tailored for particular training and infere…