8 papers
MoSE: Mixture of Slimmable Experts for Efficient and Adaptive Language Models
Nurbek Tastan, Stefanos Laskaridis, Karthik Nandakumar +1
Mixture-of-Experts (MoE) models scale large language models efficiently by sparsely activating experts, but once an expert is selected, it is executed fully. Hence, the trade-off b…
FlexRank: Nested Low-Rank Knowledge Decomposition for Adaptive Model Deployment
Riccardo Zaccone, Stefanos Laskaridis, Marco Ciccone +1
The growing scale of deep neural networks, encompassing large language models (LLMs) and vision transformers (ViTs), has made training from scratch prohibitively expensive and depl…
LoFT: Low-Rank Adaptation That Behaves Like Full Fine-Tuning
Nurbek Tastan, Stefanos Laskaridis, Martin Takac +2
Large pre-trained models are commonly adapted to downstream tasks using parameter-efficient fine-tuning methods such as Low-Rank Adaptation (LoRA), which injects small trainable lo…
NoEsis: Differentially Private Knowledge Transfer in Modular LLM Adaptation
Rob Romijnders, Stefanos Laskaridis, Ali Shahin Shamsabadi +1
Large Language Models (LLM) are typically trained on vast amounts of data from various sources. Even when designed modularly (e.g., Mixture-of-Experts), LLMs can leak privacy on th…
MELTing point: Mobile Evaluation of Language Transformers
Stefanos Laskaridis, Kleomenis Katevas, Lorenzo Minto +1
Transformers have revolutionized the machine learning landscape, gradually making their way into everyday tasks and equipping our computers with "sparks of intelligence". However,…
The Future of Consumer Edge-AI Computing
Stefanos Laskaridis, Stylianos I. Venieris, Alexandros Kouris +2
In the last decade, Deep Learning has rapidly infiltrated the consumer end, mainly thanks to hardware acceleration across devices. However, as we look towards the future, it is evi…