1 paper
Kexin Chu, Yang Zhou, Wei Zhang
Temperature-zero BF16 LLM inference is often treated as reproducible, yet the same request can emit different tokens when decoded alone or inside a larger batch. Existing fixes use…