2 papers
cs.LG2026
Unifying Sparse Attention with Hierarchical Memory for Scalable Long-Context LLM Serving
Zihan Zhao, Baotong Lu, Shengjie Lin +8
Long-context LLM serving is bottlenecked by the cost of attending over ever-growing KV caches. Dynamic sparse attention promises relief by accessing only a small, query-dependent s…
cs.AR2025
MCMComm: Hardware-Software Co-Optimization for End-to-End Communication in Multi-Chip-Modules
Ritik Raj, Shengjie Lin, William Won +1
Increasing AI computing demands and slowing transistor scaling have led to the advent of Multi-Chip-Module (MCMs) based accelerators. MCMs enable cost-effective scalability, higher…