4 citations · 4 across the 2 of their papers we have counts for
2 papers
cs.CL2026
Bielik-Minitron-7B: Compressing Large Language Models via Structured Pruning and Knowledge Distillation for the Polish Language
Remigiusz Kinas, Paweł Kiszczak, Sergio P. Perez +4
This report details the creation of Bielik-Minitron-7B, a compressed 7.35B parameter version of the Bielik-11B-v3.0 model, specifically optimized for European languages. By leverag…
cs.LG2023★ 4 cited
Training and inference of large language models using 8-bit floating point
Sergio P. Perez, Yan Zhang, James Briggs +6
FP8 formats are gaining popularity to boost the computational efficiency for training and inference of large deep learning models. Their main challenge is that a careful choice of…