paper

VUDA: Enabling Controlled Spatial Sharing of Graphics and Compute on NVIDIA GPUs

arXiv:2605.01352

Abstract

Graphics and compute increasingly share a single GPU in embodied AI simulators, AI-enabled games, and VR systems. Running these workloads concurrently can improve utilization, but contention can also compromise rendering and inference latency. Effective sharing therefore requires control over both concurrency and resource allocation. Native CUDA and Vulkan runtimes complicate this task: they place work in separate scheduling domains, while compute-oriented resource controls do not govern the graphics pipeline. We present VUDA, a system that enables controlled spatial sharing of native CUDA compute and Vulkan graphics on NVIDIA GPUs. An analysis of GPU scheduling, address translation, and workload dispatch reveals how to coordinate the two execution stacks without replacing either one. VUDA redirects CUDA channels into Vulkan's scheduling domain while preserving the runtimes' separate data address spaces. It then uses GPU front-end controls to partition execution resources between the graphics and compute pipelines, with allocations adjustable at runtime. Together, these mechanisms let applications control resource sharing without modifying GPU drivers or rewriting kernels and shaders. We evaluate VUDA across four application scenarios on three NVIDIA GPU platforms, covering throughput and latency objectives for graphics, compute, or both. In embodied AI simulation, enabling concurrency improves throughput by up to over the same asynchronous pipeline under default time sharing. In driving perception, partitioned co-execution reduces the measured inference-budget miss rate from 80.1% to zero while sustaining 60-FPS rendering.