2 papers
stat.ML2026
AIS: Adaptive Importance Sampling for Quantized RL
Jiajun Zhou, Wei Shao, Lingchao Zheng +2
Reinforcement learning (RL) for large language models (LLMs) is dominated by the cost of rollout generation, which has motivated the use of low-precision rollouts (e.g., FP8) paire…
cs.LG2026
A Case Study of Selected PTQ Baselines for Reasoning LLMs on Ascend NPU
Yuchen Luo, Fangyue Zhu, Ruining Zhou +4
Post-Training Quantization (PTQ) is crucial for efficient model deployment, yet its effectiveness on Ascend NPU remains under-explored compared to GPU architectures. This paper pre…