5 papers
WA-SpecDec: World-Aware Speculative Decoding for Vision-Language-Action Models
Zikang Wen, Yuning Zhang, Dong Yuan
Vision-language-action (VLA) policies generate robot controls autoregressively, making closed-loop latency dominated by repeated target-model forward passes. Speculative decoding r…
HorizonServe: Coordinating Request Scheduling with GPU Sharing for Omni-Model Serving
Yuning Zhang, Dong Yuan
Omni models unify text, speech, image, and multimodal reasoning in a single serving backend, but this unified deployment exposes a new scheduling problem. Requests with different o…
DuoServe-MoE: Dual-Phase Expert Prefetch and Caching for LLM Inference QoS Assurance
Yuning Zhang, Grant Pinkert, Nan Yang +2
Large Language Models (LLMs) are increasingly deployed as Internet/Web services (LLM-as-a-Service) with strict latency Service-Level Objectives (SLOs) under tight GPU memory budget…
AgentServe: Algorithm-System Co-Design for Efficient Agentic AI Serving on a Consumer-Grade GPU
Yuning Zhang, Yan Yan, Nan Yang +1
Large language models (LLMs) are increasingly deployed as AI agents that operate in short reasoning-action loops, interleaving model computation with external calls. Unlike traditi…
GuardFed: A Trustworthy Federated Learning Framework Against Dual-Facet Attacks
Yanli Li, Yanan Zhou, Zhongliang Guo +6
Federated learning (FL) enables privacy-preserving collaborative model training but remains vulnerable to adversarial behaviors that compromise model utility or fairness across sen…