5 papers
Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving
Ferran Agullo, Joan Oliveras, Chen Wang +5
Large Language Model (LLM) adapters enable low-cost model specialization, but introduce complex caching and scheduling challenges in distributed serving systems where hundreds of a…
In-Context Bias Propagation in LLM-Based Tabular Data Generation
Pol G. Recasens, Alberto Gutierrez, Jordi Torres +4
Large Language Models (LLMs) are increasingly used for synthetic tabular data generation through in-context learning (ICL), offering a practical solution for data augmentation in d…
FRIDA: Free-Rider Detection using Privacy Attacks
Pol G. Recasens, Ádám Horváth, Alberto Gutierrez-Torre +3
Federated learning is increasingly popular as it enables multiple parties with limited datasets and resources to train a machine learning model collaboratively. However, similar to…
Enabling an OpenStack-based cloud on top of RISC-V hardware
Diego Marrón, Aaron Call, Josep Ll. Berral +1
The European Union's technological sovereignty strategy centers around the RISC-V Instruction Set Architecture, with the European Processor Initiative leading efforts to build prod…
Towards Pareto Optimal Throughput in Small Language Model Serving
Pol G. Recasens, Yue Zhu, Chen Wang +5
Large language models (LLMs) have revolutionized the state-of-the-art of many different natural language processing tasks. Although serving LLMs is computationally and memory deman…