2 papers
cs.DC2026
APEX4: Efficient Pure W4A4 LLM Inference via Intra-SM Compute Rebalancing
Hong Guo, Nianhui Guo, Weixing Wang +3
W4A4 quantization promises full utilization of INT4 Tensor Cores, yet group dequantization overhead on CUDA Cores has driven existing systems to mixed-precision fallbacks. We prese…
cs.LG2026
Sample Where You Struggle: Sharpening Base Model Reasoning via Entropy-Guided Power Sampling
Hong Guo, Nianhui Guo, Christoph Meinel +1
Sampling from the sequence-level power distribution elicits RL-level reasoning from base language models without any parameter updates, but the standard Metropolis--Hastings…