2 papers
cs.IR2025
Scaling Down, Serving Fast: Compressing and Deploying Efficient LLMs for Recommendation Systems
Kayhan Behdin, Ata Fatahibaarzi, Qingquan Song +17
Large language models (LLMs) have demonstrated remarkable performance across a wide range of industrial applications, from search and recommendation systems to generative tasks. Al…
cs.CL2025
LANTERN: Scalable Distillation of Large Language Models for Job-Person Fit and Explanation
Zhoutong Fu, Yihan Cao, Yi-Lin Chen +16
Large language models (LLMs) have achieved strong performance across a wide range of natural language processing tasks. However, deploying LLMs at scale for domain specific applica…