2 papers
cs.LG2026
Reverse Distillation: Consistently Scaling Protein Language Model Representations
Darius Catrina, Christian Bepler, Samuel Sledzieski +1
Unlike the predictable scaling laws in natural language processing and computer vision, protein language models (PLMs) scale poorly: for many tasks, models within the same family p…
cs.CL2021
Distilling the Knowledge of Romanian BERTs Using Multiple Teachers
Andrei-Marius Avram, Darius Catrina, Dumitru-Clementin Cercel +4
Running large-scale pre-trained language models in computationally constrained environments remains a challenging problem yet to be addressed, while transfer learning from these mo…