8 papers
Cost-Optimal LLM Routing with Limited User Feedback under User Satisfaction Guarantees
Herbert Woisetschläger, Arastun Mammadli, Ryan Zhang +1
Inference costs for large language model (LLM) applications are rapidly growing, driven by surging demand and rising infrastructure cost. Users expect high-quality responses, and i…
Position: Let's Develop Data Probes to Fundamentally Understand How Data Affects LLM Performance
Shiqiang Wang, Herbert Woisetschläger, Hans Arno Jacobsen +1
Data is fundamental to large language models (LLMs). However, understanding of what makes certain data useful for different stages of an LLM workflow, including training, tuning, a…
Agentic Performance at the Edge: Insights from Benchmarking
Shiqiang Wang, Herbert Woisetschläger
Agentic artificial intelligence (AI) is a natural fit for Internet of Things (IoT) and edge systems, but edge deployments are often constrained to models around 8 billion parameter…
MAR-FL: A Communication Efficient Peer-to-Peer Federated Learning System
Felix Mulitze, Herbert Woisetschläger, Hans Arno Jacobsen
The convergence of next-generation wireless systems and distributed Machine Learning (ML) demands Federated Learning (FL) methods that remain efficient and robust with wireless con…
MESS+: Dynamically Learned Inference-Time LLM Routing in Model Zoos with Service Level Guarantees
Herbert Woisetschläger, Ryan Zhang, Shiqiang Wang +1
Open-weight large language model (LLM) zoos provide access to numerous high-quality models, but selecting the appropriate model for specific tasks remains challenging and requires…
Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining
Daouda Sow, Herbert Woisetschläger, Saikiran Bulusu +3
Pretraining large language models (LLMs) on vast and heterogeneous datasets is crucial for achieving state-of-the-art performance across diverse downstream tasks. However, current…