3 papers
cs.CL2025
Assortment of Attention Heads: Accelerating Federated PEFT with Head Pruning and Strategic Client Selection
Yeshwanth Venkatesha, Souvik Kundu, Priyadarshini Panda
Parameter Efficient Fine-Tuning (PEFT) has become the de-facto approach in adapting Large Language Models (LLMs) for downstream tasks in Natural Language Processing. However, its a…
cs.RO2025
Fast and Cost-effective Speculative Edge-Cloud Decoding with Early Exits
Yeshwanth Venkatesha, Souvik Kundu, Priyadarshini Panda
Large Language Models (LLMs) enable various applications on edge devices such as smartphones, wearables, and embodied robots. However, their deployment often depends on expensive c…
cs.AR2025
MEADOW: Memory-efficient Dataflow and Data Packing for Low Power Edge LLMs
Abhishek Moitra, Arkapravo Ghosh, Shrey Agarwal +3
The computational and memory challenges of large language models (LLMs) have sparked several optimization approaches towards their efficient implementation. While prior LLM-targete…