3 papers
cs.NI2026
When Scaling Fails: Network and Fabric Effects on Distributed GPU Training Performance
Dinesh Gopalan, Ratul Ali
Scaling distributed GPU training is commonly assumed to yield predictable performance gains as additional nodes are added. In practice, many large-scale deployments encounter dimin…
cs.DC2026
HQP: Sensitivity-Aware Hybrid Quantization and Pruning for Ultra-Low-Latency Edge AI Inference
Dinesh Gopalan, Ratul Ali
The escalating demand for high-fidelity, real-time inference in distributed edge-cloud environments necessitates aggressive model optimization to counteract severe latency and ener…
cs.AI2026
Scalable and Secure AI Inference in Healthcare: A Comparative Benchmarking of FastAPI and Triton Inference Server on Kubernetes
Ratul Ali
Efficient and scalable deployment of machine learning (ML) models is a prerequisite for modern production environments, particularly within regulated domains such as healthcare and…