3 papers
cs.CL2026
Scale or Reason? A Compute-Equivalent Analysis of Reasoning Distillation
Nicolas Boizard, Hippolyte Gisserot-Boukhlef, Kevin El Haddad +3
Distilling reasoning traces from strong teacher models has become the standard recipe for building capable small language models. Yet reasoning traces are 5-20 longer than…
cs.CL2026
Curate-Train-Refine: A Closed-Loop Agentic Framework for Zero Shot Classification
Gaurav Maheshwari, Kevin El Haddad
Large language models (LLMs) and high-capacity encoders have advanced zero and few-shot classification, but their inference cost and latency limit practical deployment. We propose…
cs.CL2025
Towards Cross-Tokenizer Distillation: the Universal Logit Distillation Loss for LLMs
Nicolas Boizard, Kevin El Haddad, Céline Hudelot +1
Deploying large language models (LLMs) of several billion parameters can be impractical in most industrial use cases due to constraints such as cost, latency limitations, and hardw…