1 paper · 1 filter
Ahmad Fraij, Sam Dauncey
Data scarcity drives the need for more sample-efficient large language models. In this work, we use the double descent phenomenon to holistically compare the sample efficiency of d…