activity
20242026
collaborators
Showing cs.LGShow all

7 papers · 1 filter

cs.LG2026

Super Apriel: One Checkpoint, Many Speeds

SLAM Labs, :, Oleksiy Ostapenko +13

We release Super Apriel, a 15B-parameter supernet in which every decoder layer provides four trained mixer choices -- Full Attention (FA), Sliding Window Attention (SWA), Kimi Delt…

cs.LG2026

DiffuMamba: High-Throughput Diffusion LMs with Mamba Backbone

Vaibhav Singh, Oleksiy Ostapenko, Pierre-André Noël +2

Diffusion language models (DLMs) have emerged as a promising alternative to autoregressive (AR) generation, yet their reliance on Transformer backbones limits inference efficiency…

cs.LG2025

Apriel-H1: Towards Efficient Enterprise Reasoning Models

Oleksiy Ostapenko, Luke Kumar, Raymond Li +10

Large Language Models (LLMs) achieve remarkable reasoning capabilities through transformer architectures with attention mechanisms. However, transformers suffer from quadratic time…

cs.LG2025

Unifying Autoregressive and Diffusion-Based Sequence Generation

Nima Fathi, Torsten Scholak, Pierre-André Noël

We present significant extensions to diffusion-based sequence generation models, blurring the line with autoregressive language models. We introduce hyperschedules, which assign di…

cs.LG2025

Apriel-Nemotron-15B-Thinker

Shruthan Radhakrishna, Soham Parikh, Gopal Sarda +32

While large language models (LLMs) have achieved remarkable reasoning capabilities across domains like code, math and other enterprise tasks, their significant memory and computati…

cs.LG2025

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training

Oleksiy Ostapenko, Charles Guille-Escuret, Luke Kumar +7

We introduce a framework for optimizing domain-specific dataset construction in foundation model training. Specifically, we seek a cost-efficient way to estimate the quality of dat…