Showing cs.DCShow all
2 papers · 1 filter
cs.DC2025
TokenWeave: Efficient Compute-Communication Overlap for Distributed LLM Inference
Raja Gond, Nipun Kwatra, Ramachandran Ramjee
Distributed inference of large language models (LLMs) using tensor parallelism can introduce communication overheads of % even over GPUs connected via NVLink, a high-speed GPU…
cs.DC2024
emucxl: an emulation framework for CXL-based disaggregated memory applications
Raja Gond, Purushottam Kulkarni
The emergence of CXL (Compute Express Link) promises to transform the status of interconnects between host and devices and in turn impact the design of all software layers. With it…