3 papers
cs.AI2026
Bayesian Repetition Penalty: A Principled Adjacent-Conditional Framework for Reversing Attention Collapse in Autoregressive Language Models
Wenjie Fan, Bin Ma, Dong Li
Attention collapse in autoregressive language models -- manifested as repetitive token loops where the model becomes trapped in self-reinforcing attractors -- is a persistent patho…
cs.DC2026
Distributed Generative Inference of LLM at Internet Scales with Multi-Dimensional Communication Optimization
Jiu Chen, Shuangyan Yang, Xu Xiong +4
Decentralized LLM inference distributes computation among heterogeneous nodes across the internet, offering a performant and cost-efficient solution, alternative to traditional cen…
cs.PF2025
Machine Learning-Guided Memory Optimization for DLRM Inference on Tiered Memory
Jie Ren, Bin Ma, Shuangyan Yang +4
Deep learning recommendation models (DLRMs) are widely used in industry, and their memory capacity requirements reach the terabyte scale. Tiered memory architectures provide a cost…