1 paper
Jiarui Guo, Rongle Wang, Peijun Huang +7
As LLM context windows expand and input sequences grow longer, serving systems face increasing computational and memory demands. Context parallelism (CP), which partitions the inpu…