Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD
Dongfang Li, Xiaodong Luo, Ruoyu Sun +64
Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challenges for large-scale distributed training, including severe memory pre…
cs.CL2026
LycheeDecode: Accelerating Long-Context LLM Inference via Hybrid-Head Sparse Decoding
Gang Lin, Dongfang Li, Zhuoen Chen +4
The proliferation of long-context large language models (LLMs) exposes a key bottleneck: the rapidly expanding key-value cache during decoding, which imposes heavy memory and laten…