1 paper
Abbas Zeitoun, Lucas Torroba-Hennigen, Yoon Kim
LLM architecture research generally aims to maximize model quality subject to fixed compute/latency budgets. However, many applications of interest such as edge and on-device deplo…