1 paper · 1 filter
Nathan Godey, Yoav Artzi
Training a language model suite classically requires training each model separately and serving them independently. We improve both training and inference efficiency by stacking su…