distributed cache 1hybrid sliding window attention 1kvcache optimization 1mixture-of-experts 1multimodal inference 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.AR2026
Full-Pipeline Inference Optimization for MiMo-V2.5 Series: Pushing Hybrid SWA Efficiency to the Limit
Xiaomi MiMo Team, Anqi Liu, Aoxin Ma +28
The paper describes a production-ready inference system for the MiMo-V2.5 large language model family that combines hybrid sliding window attention, sparse mixture-of-experts, and…
cs.CV2026
Diverse Normal Prototypes-Guided Contrastive Reconstruction for Medical Anomaly Detection
Luhu Li, Bin Liu, Bowen Lin +3
Anomaly detection in medical images is challenging due to limited annotations and the domain gap. Existing reconstruction-based methods often rely on frozen pre-trained encoders, r…