10 papers
GRACE: Generative Recommender Acceleration Engine for Real-Time Ads Retrieval
Zhou Fang, Yuhang Huang, Ang Zhang +12
Productionizing generative recommenders for high-volume, real-time ads retrieval creates two serving challenges: eligibility, ensuring that each generated ad is eligible for the re…
KernelEvolve: Scaling Agentic Kernel Coding for Heterogeneous AI Accelerators at Meta
Gang Liao, Hongsen Qin, Ying Wang +36
Making deep learning recommendation model (DLRM) training and inference fast and efficient is important. However, this presents three key system challenges - model architecture div…
Experience Graphs: The Data Foundation for Self-Improving Agents
Gang Liao, Yujia He, Abdullah Ozturk +22
The database community has repeatedly advanced the state of the art by recognizing that new workloads demand new system architectures. We argue that long-horizon agentic tasks -- c…
CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters
Shaoyuan Huang, Yunfeng Zhao, Na Yan +5
As Large Language Models (LLMs) are increasingly adopted in edge intelligence to power domain-specific applications and personalized services, the quality and efficiency of the LLM…
FairMT: Fairness for Heterogeneous Multi-Task Learning
Guanyu Hu, Tangzheng Lian, Na Yan +5
Fairness in machine learning has been extensively studied in single-task settings, while fair multi-task learning (MTL), especially with heterogeneous tasks (classification, detect…
Communication-Aware Knowledge Distillation for Federated LLM Fine-Tuning over Wireless Networks
Xinlu Zhang, Na Yan, Yang Su +2
Federated learning (FL) for large language models (LLMs) offers a privacy-preserving scheme, enabling clients to collaboratively fine-tune locally deployed LLMs or smaller language…