1 citations · 1 across the 3 of their papers we have counts for
3 papers
cs.IR2025★ 1 cited
Scaling Down, Serving Fast: Compressing and Deploying Efficient LLMs for Recommendation Systems
Kayhan Behdin, Ata Fatahibaarzi, Qingquan Song +17
Large language models (LLMs) have demonstrated remarkable performance across a wide range of industrial applications, from search and recommendation systems to generative tasks. Al…
cs.DC2025
Locality-aware Fair Scheduling in LLM Serving
Shiyi Cao, Yichuan Wang, Ziming Mao +10
Large language model (LLM) inference workload dominates a wide variety of modern AI applications, ranging from multi-turn conversation to document analysis. Balancing fairness and…
cs.CR2023
NFT.mine: An xDeepFM-based Recommender System for Non-fungible Token (NFT) Buyers
Shuwei Li, Yucheng Jin, Pin-Lun Hsu +1
Non-fungible token (NFT) is a tradable unit of data stored on the blockchain which can be associated with some digital asset as a certification of ownership. The past several years…