2 papers
cs.CL2026
Less Experts, Faster Decoding: Cost-Aware Speculative Decoding for Mixture-of-Experts
Jincheng Xie, Runheng Liu, Heyan Huang +4
Sparse Mixture-of-Experts (MoE) models have become an important approach for scaling Large Language Models (LLMs), but their inference efficiency depends strongly on expert activat…
cs.AI2025
Lita: Light Agent Uncovers the Agentic Coding Capabilities of LLMs
Hankun Dai, Maoquan Wang, Mengnan Qi +6
Large language models (LLMs) are increasingly being applied to programming tasks, ranging from single-turn code completion to autonomous agents. Current code agent designs frequent…