3 papers
cs.IR2025
Scaling Down, Serving Fast: Compressing and Deploying Efficient LLMs for Recommendation Systems
Kayhan Behdin, Ata Fatahibaarzi, Qingquan Song +17
Large language models (LLMs) have demonstrated remarkable performance across a wide range of industrial applications, from search and recommendation systems to generative tasks. Al…
cs.DS2025
LLM Query Scheduling with Prefix Reuse and Latency Constraints
Gregory Dexter, Shao Tang, Ata Fatahi Baarzi +3
The efficient deployment of large language models (LLMs) in online settings requires optimizing inference performance under stringent latency constraints, particularly the time-to-…
cs.DS2024
The Space Complexity of Approximating Logistic Loss
Gregory Dexter, Petros Drineas, Rajiv Khanna
We provide space complexity lower bounds for data structures that approximate logistic loss up to -relative error on a logistic regression problem with data $\mathbf{X} \in \mat…