4 papers
Compute-Bounded Security Assurance - Coverage, Verification, and Response under Resource Constraints
Jithin VG, Ditto PS
Additional inference compute can increase the number of correctly resolved security-assurance tasks, but repeated success, unique coverage, accepted evidence, and operational prote…
GPU-Virt-Bench: A Comprehensive Benchmarking Framework for Software-Based GPU Virtualization Systems
Jithin VG, Ditto PS
The proliferation of GPU-accelerated workloads, particularly in artificial intelligence and large language model (LLM) inference, has created unprecedented demand for efficient GPU…
Efficient Hybrid Inference for LLMs: Reward-Based Token Modelling with Selective Cloud Assistance
Adarsh MS, Jithin VG, Ditto PS
Large language models (LLMs) are known for their exceptional performance across a range of natural language processing tasks, but their deployment comes at a high computational and…
Inference Acceleration for Large Language Models on CPUs
Ditto PS, Jithin VG, Adarsh MS
In recent years, large language models have demonstrated remarkable performance across various natural language processing (NLP) tasks. However, deploying these models for real-wor…