Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
WARP: Guaranteed Inner-Layer Repair of NLP Transformers
Hsin-Ling Hsu, Min-Yu Chen, Nai-Chia Chen +3
Transformer-based NLP models remain vulnerable to adversarial perturbations, yet existing repair methods face a fundamental trade-off: gradient-based approaches offer flexibility b…
cs.LG2025
Muon is Scalable for LLM Training
Jingyuan Liu, Jianlin Su, Xingcheng Yao +25
Recently, the Muon optimizer based on matrix orthogonalization has demonstrated strong results in training small-scale language models, but the scalability to larger models has not…
cs.LG2025
MoBA: Mixture of Block Attention for Long-Context LLMs
Enzhe Lu, Zhejun Jiang, Jingyuan Liu +22
Scaling the effective context length is essential for advancing large language models (LLMs) toward artificial general intelligence (AGI). However, the quadratic increase in comput…