1 paper
DatologyAI, :, Matthew L. Leavitt +8
Inference efficiency is typically pursued by shrinking the model: distillation, pruning, quantization, and sparse routing each lower per-token cost while treating token count as fi…