collaborators

7 papers

cs.DC2026

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs

Zhihao Xu, Hao Zhong, Zeting Zhou +9

This paper aims to enable computation- and communication-efficient GPU sharing across devices within local area networks (LANs), facilitating ubiquitous AI inference on heterogeneo…

cs.CV2026

CaT-GS: Efficient 3DGS Rendering for Large Scale Scenes via Inter-frame Caching and Tile Scheduling

Tingjia Zhang, Bo Chen, Shengzhong Liu +2

Recent breakthroughs in 3D Gaussian Splatting (3DGS) have advanced neural rendering with high fidelity and speed. However, its performance degrades significantly in large-scale sce…

cs.LG2025

TokenFlow: Responsive LLM Text Streaming Serving under Request Burst via Preemptive Scheduling

Junyi Chen, Chuheng Du, Renyuan Liu +6

Real-time LLM interactions demand streamed token generations, where text tokens are progressively generated and delivered to users while balancing two objectives: responsiveness (i…

cs.OS2025

Nova: Real-Time Agentic Vision-Language Model Serving with Adaptive Cross-Stage Parallelization

Yuhang Xu, Shengzhong Liu, Dong Zhang +3

This paper presents Nova, a real-time scheduling framework for serving agentic vision-language models (VLMs) on a single GPU with balanced per-request latency and overall request p…

cs.LG2025

ConfClip: Confidence-Weighted and Clipped Reward for Reinforcement Learning in LLMs

Bonan Zhang, Zhongqi Chen, Bowen Song +3

Reinforcement learning (RL) has become a standard paradigm for refining large language models (LLMs) beyond pre-training and instruction tuning. A prominent line of work is RL with…

cs.CV2025

Responsive DNN Adaptation for Video Analytics against Environment Shift via Hierarchical Mobile-Cloud Collaborations

Maozhe Zhao, Shengzhong Liu, Fan Wu +1

Mobile video analysis systems often encounter various deploying environments, where environment shifts present greater demands for responsiveness in adaptations of deployed "expert…