2 papers
cs.LG2025
RLHFSpec: Breaking the Efficiency Bottleneck in RLHF Training via Adaptive Drafting
Siqi Wang, Hailong Yang, Junjie Zhu +3
Reinforcement Learning from Human Feedback (RLHF) is an important fine-tuning technique for large language models (LLMs) and comprises three stages: generation, inference, and trai…
cs.DC2025
LOw-cOst yet High-Performant Sparse Matrix-Matrix Multiplication on Arm SME Architectures
Kelun Lei, Hailong Yang, Kaige Zhang +8
Sparse matrix-dense matrix multiplication (SpMM) is a critical kernel in both scientific computing and emerging graph learning workloads. The recent Armv9 architecture introduces S…