4 papers · 1 filter
Cost-Optimal LLM Routing with Limited User Feedback under User Satisfaction Guarantees
Herbert Woisetschläger, Arastun Mammadli, Ryan Zhang +1
Inference costs for large language model (LLM) applications are rapidly growing, driven by surging demand and rising infrastructure cost. Users expect high-quality responses, and i…
Clin-JEPA: A Multi-Phase Co-Training Framework for Joint-Embedding Predictive Pretraining on EHR Patient Trajectories
Yixuan Yang, Mehak Arora, Ryan Zhang +10
We present Clin-JEPA, a multi-phase co-training framework for joint-embedding predictive (JEPA) pretraining on EHR patient trajectories. JEPA architectures have enabled latent-spac…
MESS+: Dynamically Learned Inference-Time LLM Routing in Model Zoos with Service Level Guarantees
Herbert Woisetschläger, Ryan Zhang, Shiqiang Wang +1
Open-weight large language model (LLM) zoos provide access to numerous high-quality models, but selecting the appropriate model for specific tasks remains challenging and requires…
MESS+: Energy-Optimal Inferencing in Language Model Zoos with Service Level Guarantees
Ryan Zhang, Herbert Woisetschläger, Shiqiang Wang +1
Open-weight large language model (LLM) zoos allow users to quickly integrate state-of-the-art models into systems. Despite increasing availability, selecting the most appropriate m…