2 papers
cs.LG2026
Breaking Model Lock-in: Cost-Efficient Zero-Shot LLM Routing via a Universal Latent Space
Cheng Yan, Wuyang Zhang, Zhiyuan Ning +5
The rapid proliferation of Large Language Models (LLMs) has led to a fragmented and inefficient ecosystem, a state of ``model lock-in'' where seamlessly integrating novel models re…
cs.OS2025
UrgenGo: Urgency-Aware Transparent GPU Kernel Launching for Autonomous Driving
Hanqi Zhu, Wuyang Zhang, Xinran Zhang +5
The rapid advancements in autonomous driving have introduced increasingly complex, real-time GPU-bound tasks critical for reliable vehicle operation. However, the proprietary natur…