2 papers
cs.LG2026
MXAttention: Data-Free Optimal Scaling and Pre-Normalization Quantization for MXFP4 Attention
Jianlin Yu, Jing Lin, Linghui Kong +13
The quadratic cost of attention is a major bottleneck in diffusion-based video generation models. MXFP4 attention provides a promising path toward efficient inference, but direct M…
cs.LG2026
ACRL: Adaptive Control of Training-Inference Discrepancy for Stable Reinforcement Learning
Wenwu Fan, Qihong Lin, Zhijie Xia +4
Reinforcement Learning (RL) training for Large Language Models (LLMs) often suffers from instability due to the discrepancy between training and inference. This training-inference…