5 citations · 12 across the 6 of their papers we have counts for
Showing cs.DCShow all
3 papers · 1 filter
cs.DC2026
WeightBridge: An Efficient Weight Transfer Library for Reinforcement Learning
Xuanlin Jiang, Samuel Hsia, Michael Kuchnik +3
Weight transfer - the propagation of updated parameters from trainers to rollout generators - is becoming an important performance bottleneck in reinforcement learning (RL) systems…
cs.DC2024★ 1 cited
NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference
Xuanlin Jiang, Yang Zhou, Shiyi Cao +2
Online LLM inference powers many exciting applications such as intelligent chatbots and autonomous agents. Modern LLM inference engines widely rely on request batching to improve i…
cs.DC2024★ 5 cited
RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation
Chao Jin, Zili Zhang, Xuanlin Jiang +4
Retrieval-Augmented Generation (RAG) has shown significant improvements in various natural language processing tasks by integrating the strengths of large language models (LLMs) an…