3 papers
cs.LG2026
Advancing Model Refinement: Muon-Optimized Distillation and Quantization for LLM Deployment
Jacob Sander, Brian Jalaian, Venkat R. Dasari
Large Language Models (LLMs) enable advanced natural language processing but face deployment challenges on resource-constrained edge devices due to high computational, memory, and…
cs.CV2025
GFT: Graph Feature Tuning for Efficient Point Cloud Analysis
Manish Dhakal, Venkat R. Dasari, Rajshekhar Sunderraman +1
Parameter-efficient fine-tuning (PEFT) significantly reduces computational and memory costs by updating only a small subset of the model's parameters, enabling faster adaptation to…
cs.LG2025
Constrained Edge AI Deployment: Fine-Tuning vs Distillation for LLM Compression
Jacob Sander, David Moe, Achraf Cohen +3
Modern foundational models are often compressed via a combination of structured pruning and re-training to meet the strict compute, memory, and connectivity constraints of edge dep…