Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Mixture of Universal Experts: Scaling Virtual Width via Depth-Width Transformation
Yilong Chen, Naibin Gu, Junyuan Shang +8
Mixture-of-Experts (MoE) decouples model capacity from per-token computation, yet their scalability remains limited by the physical dimensions of depth and width. To overcome this,…
cs.LG2025
Semantic Energy: Detecting LLM Hallucination Beyond Entropy
Huan Ma, Jiadong Pan, Jing Liu +7
Large Language Models (LLMs) are being increasingly deployed in real-world applications, but they remain susceptible to hallucinations, which produce fluent yet incorrect responses…