3 papers
cs.LG2026
SOLARIS: Speculative Offloading of Latent-bAsed Representation for Inference Scaling
Zikun Liu, Liang Luo, Qianru Li +31
Recent advances in recommendation scaling laws have led to foundation models of unprecedented complexity. While these models offer superior performance, their computational demands…
cs.DC2025
SpecInF: Exploiting Idle GPU Resources in Distributed DL Training via Speculative Inference Filling
Cunchi Lv, Xiao Shi, Dong Liang +2
Deep Learning (DL), especially with Large Language Models (LLMs), brings benefits to various areas. However, DL training systems usually yield prominent idling GPU resources due to…
cs.IR2024
ERCache: An Efficient and Reliable Caching Framework for Large-Scale User Representations in Meta's Ads System
Fang Zhou, Yaning Huang, Dong Liang +21
The increasing complexity of deep learning models used for calculating user representations presents significant challenges, particularly with limited computational resources and s…