2 citations · 2 across the 4 of their papers we have counts for
3 papers · 1 filter
Advancing Model Refinement: Muon-Optimized Distillation and Quantization for LLM Deployment
Jacob Sander, Brian Jalaian, Venkat R. Dasari
Large Language Models (LLMs) enable advanced natural language processing but face deployment challenges on resource-constrained edge devices due to high computational, memory, and…
Constrained Edge AI Deployment: Fine-Tuning vs Distillation for LLM Compression
Jacob Sander, David Moe, Achraf Cohen +3
Modern foundational models are often compressed via a combination of structured pruning and re-training to meet the strict compute, memory, and connectivity constraints of edge dep…
On Accelerating Edge AI: Optimizing Resource-Constrained Environments
Jacob Sander, Achraf Cohen, Venkat R. Dasari +2
Resource-constrained edge deployments demand AI solutions that balance high performance with stringent compute, memory, and energy limitations. In this survey, we present a compreh…