2 papers
cs.DC2026
Efficient, VRAM-Constrained xLM Inference on Clients
Aditya Ukarande, Deep Shekhar, Marc Blackstein +1
To usher in the next round of client AI innovation, there is an urgent need to enable efficient, lossless inference of high-accuracy large language models (LLMs) and vision languag…
cs.DC2025
RIMMS: Runtime Integrated Memory Management System for Heterogeneous Computing
Serhan Gener, Aditya Ukarande, Shilpa Mysore Srinivasa Murthy +5
Efficient memory management in heterogeneous systems is increasingly challenging due to diverse compute architectures (e.g., CPU, GPU, FPGA) and dynamic task mappings not known at…