4 papers
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs
Zhihao Xu, Hao Zhong, Zeting Zhou +9
This paper aims to enable computation- and communication-efficient GPU sharing across devices within local area networks (LANs), facilitating ubiquitous AI inference on heterogeneo…
BubbleSpec: Turning Long-Tail Bubbles into Speculative Rollout Drafts for Synchronous Reinforcement Learning
Yuhang Xu, Kaibin Tian, Yang Tian +6
Reinforcement Learning (RL) has become a cornerstone for improving the performance of Large Language Models (LLMs). However, its rollout phase constitutes a significant efficiency…
Nova: Real-Time Agentic Vision-Language Model Serving with Adaptive Cross-Stage Parallelization
Yuhang Xu, Shengzhong Liu, Dong Zhang +3
This paper presents Nova, a real-time scheduling framework for serving agentic vision-language models (VLMs) on a single GPU with balanced per-request latency and overall request p…
Intersection-free Robot Manipulation with Soft-Rigid Coupled Incremental Potential Contact
Wenxin Du, Siqiong Yao, Xinlei Wang +3
This paper presents a novel simulation platform, ZeMa, designed for robotic manipulation tasks concerning soft objects. Such simulation ideally requires three properties: two-way s…