5 papers
Scale or Reason? A Compute-Equivalent Analysis of Reasoning Distillation
Nicolas Boizard, Hippolyte Gisserot-Boukhlef, Kevin El Haddad +3
Distilling reasoning traces from strong teacher models has become the standard recipe for building capable small language models. Yet reasoning traces are 5-20 longer than…
Curate-Train-Refine: A Closed-Loop Agentic Framework for Zero Shot Classification
Gaurav Maheshwari, Kevin El Haddad
Large language models (LLMs) and high-capacity encoders have advanced zero and few-shot classification, but their inference cost and latency limit practical deployment. We propose…
Towards Cross-Tokenizer Distillation: the Universal Logit Distillation Loss for LLMs
Nicolas Boizard, Kevin El Haddad, Céline Hudelot +1
Deploying large language models (LLMs) of several billion parameters can be impractical in most industrial use cases due to constraints such as cost, latency limitations, and hardw…
ASR Benchmarking: Need for a More Representative Conversational Dataset
Gaurav Maheshwari, Dmitry Ivanov, Théo Johannet +1
Automatic Speech Recognition (ASR) systems have achieved remarkable performance on widely used benchmarks such as LibriSpeech and Fleurs. However, these benchmarks do not adequatel…
Efficacy of Synthetic Data as a Benchmark
Gaurav Maheshwari, Dmitry Ivanov, Kevin El Haddad
Large language models (LLMs) have enabled a range of applications in zero-shot and few-shot learning settings, including the generation of synthetic datasets for training and testi…