Showing 2025Show all
2 papers · 1 filter
cs.LG2025
AMLA: MUL by ADD in FlashAttention Rescaling
Qichen Liao, Chengqiu Hu, Fangzheng Miao +8
Multi-head Latent Attention (MLA) significantly reduces KVCache memory usage in Large Language Models while introducing substantial computational overhead and intermediate variable…
cs.CV2025
Dynamic Pattern Alignment Learning for Pretraining Lightweight Human-Centric Vision Models
Xuanhan Wang, Huimin Deng, Ke Liu +3
Human-centric vision models (HVMs) have achieved remarkable generalization due to large-scale pretraining on massive person images. However, their dependence on large neural archit…