1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.IR2025
SilverTorch: A Unified Model-based System to Democratize Large-Scale Recommendation on GPUs
Bi Xue, Hong Wu, Lei Chen +29
Serving deep learning based recommendation models (DLRM) at scale is challenging. Existing approaches rely on dedicated ANN indexing and filtering services on CPUs, suffering from…
cs.DC2024★ 1 cited
ENOVA: Autoscaling towards Cost-effective and Stable Serverless LLM Serving
Tao Huang, Pengfei Chen, Kyoka Gong +5
Since the increasing popularity of large language model (LLM) backend systems, it is common and necessary to deploy stable serverless serving of LLM on multi-GPU clusters with auto…