5 papers
Making Every Verified Token Count: Adaptive Verification for MoE Speculative Decoding
Lehan Pan, Ziyang Tao, Xiao Wang +3
Tree-based speculative decoding accelerates autoregressive generation by verifying multiple draft candidates in parallel, but this advantage weakens for sparse Mixture-of-Experts (…
EvoDiagram: Agentic Editable Diagram Creation via Design Expertise Evolution
Tianfu Wang, Leilei Ding, Ziyang Tao +13
High-fidelity diagram creation requires the complex orchestration of semantic topology, visual styling, and spatial layout, posing a significant challenge for automated systems. Ex…
Breaking Model Lock-in: Cost-Efficient Zero-Shot LLM Routing via a Universal Latent Space
Cheng Yan, Wuyang Zhang, Zhiyuan Ning +5
The rapid proliferation of Large Language Models (LLMs) has led to a fragmented and inefficient ecosystem, a state of ``model lock-in'' where seamlessly integrating novel models re…
GRACE-MoE: Grouping and Replication with Locality-Aware Routing for Efficient Distributed MoE Inference
Yu Han, Lehan Pan, Jie Peng +4
Sparse Mixture of Experts (SMoE) enables scalable parameter growth in large language models (LLMs) by selectively activating a subset of experts, and its large parameter count nece…
UrgenGo: Urgency-Aware Transparent GPU Kernel Launching for Autonomous Driving
Hanqi Zhu, Wuyang Zhang, Xinran Zhang +5
The rapid advancements in autonomous driving have introduced increasingly complex, real-time GPU-bound tasks critical for reliable vehicle operation. However, the proprietary natur…