4 papers
Robust Batch-Level Query Routing for Large Language Models under Cost and Capacity Constraints
Jelena Markovic-Voronov, Kayhan Behdin, Yuanda Xu +3
We study the problem of routing queries to large language models (LLMs) under cost, GPU resources, and concurrency constraints. Prior per-query routing methods often fail to contro…
DuaLip-GPU Technical Report
Gregory Dexter, Aida Rahmattalabi, Sanjana Garg +6
Large-scale linear programs (LPs) arise in many decision systems, including ranking, allocation, and matching problems that must be solved repeatedly at massive scale. Prior work s…
Scaling Down, Serving Fast: Compressing and Deploying Efficient LLMs for Recommendation Systems
Kayhan Behdin, Ata Fatahibaarzi, Qingquan Song +17
Large language models (LLMs) have demonstrated remarkable performance across a wide range of industrial applications, from search and recommendation systems to generative tasks. Al…
Efficient user history modeling with amortized inference for deep learning recommendation models
Lars Hertel, Neil Daftary, Fedor Borisyuk +2
We study user history modeling via Transformer encoders in deep learning recommendation models (DLRM). Such architectures can significantly improve recommendation quality, but usua…