activity
20222026
most citedFew-Bit Backward: Quantized Gradients of Activation Functions for Memory Footprint Reduction

5 citations · 11 across the 9 of their papers we have counts for

collaborators

9 papers

cs.CL2026

Shieldstral

Antonia Calvi, Avinash Sooriyarachchi, Giada Pistilli +273

We introduce Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms models nearly 7 its size on text safety benchmarks and set…

cs.AI2026

Voxtral Realtime

Mistral-AI, :, Alexander H. Liu +166

We introduce Voxtral Realtime, a natively streaming automatic speech recognition model that matches offline transcription quality at sub-second latency. Unlike approaches that adap…

cs.CL20261 cited

Ministral 3

Alexander H. Liu, Kartik Khandelwal, Sandeep Subramanian +116

We introduce the Ministral 3 series, a family of parameter-efficient dense language models designed for compute and memory constrained applications, available in three model sizes:…

cs.SE20252 cited

Devstral: Fine-tuning Language Models for Coding Agent Applications

Abhinav Rastogi, Adam Yang, Albert Q. Jiang +100

We introduce Devstral-Small, a lightweight open source model for code agents with the best performance among models below 100B size. In this technical report, we give an overview o…

cs.CV2024

Tensor-Train Point Cloud Compression and Efficient Approximate Nearest-Neighbor Search

Georgii Novikov, Alexander Gneushev, Alexey Kadeishvili +1

Nearest-neighbor search in large vector databases is crucial for various machine learning applications. This paper introduces a novel method using tensor-train (TT) low-rank tensor…

cs.LG2024

Inverted Activations: Reducing Memory Footprint in Neural Network Training

Georgii Novikov, Ivan Oseledets

The scaling of neural networks with increasing data and model sizes necessitates the development of more efficient deep learning algorithms. A significant challenge in neural netwo…