7 papers · 1 filter
Super Apriel: One Checkpoint, Many Speeds
SLAM Labs, :, Oleksiy Ostapenko +13
We release Super Apriel, a 15B-parameter supernet in which every decoder layer provides four trained mixer choices -- Full Attention (FA), Sliding Window Attention (SWA), Kimi Delt…
DiffuMamba: High-Throughput Diffusion LMs with Mamba Backbone
Vaibhav Singh, Oleksiy Ostapenko, Pierre-André Noël +2
Diffusion language models (DLMs) have emerged as a promising alternative to autoregressive (AR) generation, yet their reliance on Transformer backbones limits inference efficiency…
Apriel-H1: Towards Efficient Enterprise Reasoning Models
Oleksiy Ostapenko, Luke Kumar, Raymond Li +10
Large Language Models (LLMs) achieve remarkable reasoning capabilities through transformer architectures with attention mechanisms. However, transformers suffer from quadratic time…
Unifying Autoregressive and Diffusion-Based Sequence Generation
Nima Fathi, Torsten Scholak, Pierre-André Noël
We present significant extensions to diffusion-based sequence generation models, blurring the line with autoregressive language models. We introduce hyperschedules, which assign di…
Apriel-Nemotron-15B-Thinker
Shruthan Radhakrishna, Soham Parikh, Gopal Sarda +32
While large language models (LLMs) have achieved remarkable reasoning capabilities across domains like code, math and other enterprise tasks, their significant memory and computati…
Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training
Oleksiy Ostapenko, Charles Guille-Escuret, Luke Kumar +7
We introduce a framework for optimizing domain-specific dataset construction in foundation model training. Specifically, we seek a cost-efficient way to estimate the quality of dat…