collaborators

7 papers

cs.DC2026

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving

Pol G. Recasens, Ferran Agullo, Yue Zhu +3

Large language model (LLM) serving commonly increases batch size to improve throughput, but performance eventually reaches a deployment-dependent plateau beyond which larger batche…

cs.DC2026

Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving

Ferran Agullo, Joan Oliveras, Chen Wang +5

Large Language Model (LLM) adapters enable low-cost model specialization, but introduce complex caching and scheduling challenges in distributed serving systems where hundreds of a…

cs.LG2026

In-Context Bias Propagation in LLM-Based Tabular Data Generation

Pol G. Recasens, Alberto Gutierrez, Jordi Torres +4

Large Language Models (LLMs) are increasingly used for synthetic tabular data generation through in-context learning (ICL), offering a practical solution for data augmentation in d…

cs.PF2025

A Data-driven ML Approach for Maximizing Performance in LLM-Adapter Serving

Ferran Agullo, Joan Oliveras, Chen Wang +5

With the rapid adoption of Large Language Models (LLMs), LLM-adapters have become increasingly common, providing lightweight specialization of large-scale models. Serving hundreds…

cs.LG2025

FRIDA: Free-Rider Detection using Privacy Attacks

Pol G. Recasens, Ádám Horváth, Alberto Gutierrez-Torre +3

Federated learning is increasingly popular as it enables multiple parties with limited datasets and resources to train a machine learning model collaboratively. However, similar to…

cs.CL2025

Towards Pareto Optimal Throughput in Small Language Model Serving

Pol G. Recasens, Yue Zhu, Chen Wang +5

Large language models (LLMs) have revolutionized the state-of-the-art of many different natural language processing tasks. Although serving LLMs is computationally and memory deman…