2 papers
cs.LG2026
Evaluating CUDA Tile for AI Workloads on Hopper and Blackwell GPUs
Divakar Kumar Yadav, Tian Zhao, Deepak Kumar
NVIDIA's CUDA Tile (CuTile) introduces a Python-based, tile-centric abstraction for GPU kernel development that aims to simplify programming while retaining Tensor Core and Tensor…
cs.LG2026
Hybrid JIT-CUDA Graph Optimization for Low-Latency Large Language Model Inference
Divakar Kumar Yadav, Tian Zhao
Large Language Models (LLMs) have achieved strong performance across natural language and multimodal tasks, yet their practical deployment remains constrained by inference latency…