1 paper · 1 filter
Vidya Srinivas, Malek Itani, Tuochao Chen +3
Models for low-latency, streaming applications could benefit from the knowledge capacity of larger models, but edge devices cannot run these models due to resource constraints. A p…