5 papers
OTPrune: Distribution-Aligned Visual Token Pruning via Optimal Transport
Xiwen Chen, Wenhui Zhu, Gen Li +9
Multi-modal large language models (MLLMs) achieve strong visual-language reasoning but suffer from high inference cost due to redundant visual tokens. Recent work explores visual t…
MixLM: High-Throughput and Effective LLM Ranking via Text-Embedding Mix-Interaction
Guoyao Li, Ran He, Shusen Jing +21
Large language models (LLMs) excel at capturing semantic nuances and therefore show impressive relevance ranking performance in modern recommendation and search systems. However, t…
Scaling Up Efficient Small Language Models Serving and Deployment for Semantic Job Search
Kayhan Behdin, Qingquan Song, Sriram Vasudevan +17
Large Language Models (LLMs) have demonstrated impressive quality when applied to predictive tasks such as relevance ranking and semantic search. However, deployment of such LLMs r…
LiNR: Model Based Neural Retrieval on GPUs at LinkedIn
Fedor Borisyuk, Qingquan Song, Mingzhou Zhou +11
This paper introduces LiNR, LinkedIn's large-scale, GPU-based retrieval system. LiNR supports a billion-sized index on GPU models. We discuss our experiences and challenges in crea…
LiRank: Industrial Large Scale Ranking Models at LinkedIn
Fedor Borisyuk, Mingzhou Zhou, Qingquan Song +31
We present LiRank, a large-scale ranking framework at LinkedIn that brings to production state-of-the-art modeling architectures and optimization methods. We unveil several modelin…