Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
REFLEX: Rethinking MoE Inference as Refinement-Aware Compute Allocation in Diffusion Language Models
Xiang Xia, Cheng Yan, Yiming Zhang +3
Mixture-of-experts (MoE) models increase parameter capacity by activating only a small subset of experts for each token. This conditional-computation paradigm has enabled autoregre…
cs.AI2026
UPAIR: Diagnosing Reasoning States via Uncertainty-Progress Alignment for Selective Intervention
Cheng Yan, Guangyang Ye, Wuyang Zhang +5
While test-time scaling improves the problem-solving ability of large reasoning models (LRMs) through additional inference-time computation, it can also exacerbate overthinking and…