1 paper · 1 filter
Andy Zeyi Liu, Elliot Paquette, John Sous
Training loss and throughput can hide distinct internal representation in language-model training. To examine these hidden mechanics, we use spectral measurements as practical and…