6 papers
MixLM: High-Throughput and Effective LLM Ranking via Text-Embedding Mix-Interaction
Guoyao Li, Ran He, Shusen Jing +21
Large language models (LLMs) excel at capturing semantic nuances and therefore show impressive relevance ranking performance in modern recommendation and search systems. However, t…
Scaling Up Efficient Small Language Models Serving and Deployment for Semantic Job Search
Kayhan Behdin, Qingquan Song, Sriram Vasudevan +17
Large Language Models (LLMs) have demonstrated impressive quality when applied to predictive tasks such as relevance ranking and semantic search. However, deployment of such LLMs r…
Generative Sequential Notification Optimization via Multi-Objective Decision Transformers
Borja Ocejo, Ruofan Wang, Ke Liu +7
Notifications are an important communication channel for delivering timely and relevant information. Optimizing their delivery involves addressing complex sequential decision-makin…
Large Scalable Cross-Domain Graph Neural Networks for Personalized Notification at LinkedIn
Shihai He, Julie Choi, Tianqi Li +8
Notification recommendation systems are critical to driving user engagement on professional platforms like LinkedIn. Designing such systems involves integrating heterogeneous signa…
Scaling Down, Serving Fast: Compressing and Deploying Efficient LLMs for Recommendation Systems
Kayhan Behdin, Ata Fatahibaarzi, Qingquan Song +17
Large language models (LLMs) have demonstrated remarkable performance across a wide range of industrial applications, from search and recommendation systems to generative tasks. Al…
Efficient user history modeling with amortized inference for deep learning recommendation models
Lars Hertel, Neil Daftary, Fedor Borisyuk +2
We study user history modeling via Transformer encoders in deep learning recommendation models (DLRM). Such architectures can significantly improve recommendation quality, but usua…