6 papers
Large-Scale Regularized Matching on GPU Clusters
Aida Rahmattalabi, Gregory Dexter, Sanjana Garg +5
Production decision systems such as ad allocation or content matching involve millions of users and thousands of items, reducing to large-scale linear programs with sparse block-di…
Robust Batch-Level Query Routing for Large Language Models under Cost and Capacity Constraints
Jelena Markovic-Voronov, Kayhan Behdin, Yuanda Xu +3
We study the problem of routing queries to large language models (LLMs) under cost, GPU resources, and concurrency constraints. Prior per-query routing methods often fail to contro…
DuaLip-GPU Technical Report
Gregory Dexter, Aida Rahmattalabi, Sanjana Garg +6
Large-scale linear programs (LPs) arise in many decision systems, including ranking, allocation, and matching problems that must be solved repeatedly at massive scale. Prior work s…
Scaling Down, Serving Fast: Compressing and Deploying Efficient LLMs for Recommendation Systems
Kayhan Behdin, Ata Fatahibaarzi, Qingquan Song +17
Large language models (LLMs) have demonstrated remarkable performance across a wide range of industrial applications, from search and recommendation systems to generative tasks. Al…
End-to-end Feature Selection Approach for Learning Skinny Trees
Shibal Ibrahim, Kayhan Behdin, Rahul Mazumder
We propose a new optimization-based approach for feature selection in tree ensembles, an important problem in statistics and machine learning. Popular tree ensemble toolkits e.g.,…
Efficient user history modeling with amortized inference for deep learning recommendation models
Lars Hertel, Neil Daftary, Fedor Borisyuk +2
We study user history modeling via Transformer encoders in deep learning recommendation models (DLRM). Such architectures can significantly improve recommendation quality, but usua…