activity
20242026
collaborators

6 papers

cs.LG2026

From Features to Transformers: Redefining Ranking for Scalable Impact

Fedor Borisyuk, Lars Hertel, Ganesh Parameswaran +14

We present LiGR, a large-scale ranking framework developed at LinkedIn that brings state-of-the-art transformer-based modeling architectures into production. We introduce a modifie…

cs.DS2026

LLM Query Scheduling with Prefix Reuse and Latency Constraints

Gregory Dexter, Shao Tang, Ata Fatahi Baarzi +3

The efficient deployment of large language models (LLMs) in online settings requires optimizing inference performance under stringent latency constraints, particularly the time-to-…

cs.IR2025

Scaling Down, Serving Fast: Compressing and Deploying Efficient LLMs for Recommendation Systems

Kayhan Behdin, Ata Fatahibaarzi, Qingquan Song +17

Large language models (LLMs) have demonstrated remarkable performance across a wide range of industrial applications, from search and recommendation systems to generative tasks. Al…

cs.CL2025

LANTERN: Scalable Distillation of Large Language Models for Job-Person Fit and Explanation

Zhoutong Fu, Yihan Cao, Yi-Lin Chen +16

Large language models (LLMs) have achieved strong performance across a wide range of natural language processing tasks. However, deploying LLMs at scale for domain specific applica…

cs.IR2025

360Brew: A Decoder-only Foundation Model for Personalized Ranking and Recommendation

Hamed Firooz, Maziar Sanjabi, Adrian Englhardt +20

Ranking and recommendation systems are the foundation for numerous online experiences, ranging from search results to personalized content delivery. These systems have evolved into…

cs.LG2024

Efficient user history modeling with amortized inference for deep learning recommendation models

Lars Hertel, Neil Daftary, Fedor Borisyuk +2

We study user history modeling via Transformer encoders in deep learning recommendation models (DLRM). Such architectures can significantly improve recommendation quality, but usua…