3 papers
cs.LG2026
Quad Length Codes for Lossless Compression of e4m3
Aditya Agrawal, Albert Magyar, Hiteshwar Eswaraiah +5
Training and serving Large Language Models (LLMs) relies heavily on parallelization and collective operations, which are frequently bottlenecked by network bandwidth. Lossless comp…
cs.LG2026
Single-Stage Huffman Encoder for ML Compression
Aditya Agrawal, Albert Magyar, Hiteshwar Eswaraiah +5
Training and serving Large Language Models (LLMs) require partitioning data across multiple accelerators, where collective operations are frequently bottlenecked by network bandwid…
cs.AR2025
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference
Yufeng Gu, Alireza Khadem, Sumanth Umesh +5
Large Language Model (LLM) inference uses an autoregressive manner to generate one token at a time, which exhibits notably lower operational intensity compared to earlier Machine L…