2 papers
cs.CL2026
Accelerating Large Language Model Inference with Self-Supervised Early Exits
Florian Valade
This paper presents a modular approach to accelerate inference in large language models (LLMs) by adding early exit heads at intermediate transformer layers. Each head is trained i…
stat.ML2025
EERO: Early Exit with Reject Option for Efficient Classification with limited budget
Florian Valade, Mohamed Hebiri, Paul Gay
The increasing complexity of advanced machine learning models requires innovative approaches to manage computational resources effectively. One such method is the Early Exit strate…