3 papers
cs.DC2026
Fine-Tuning and Serving Gemma 4 31B on Google Cloud TPU: A Technical Comparison with GPU Baselines
Jatin Kishnani, Mayank Goel, Amit Singh +2
We present the first end-to-end demonstration of fine-tuning and serving Google's Gemma 4 31B model on TPU hardware, providing an empirical comparison of TPU and GPU platforms for…
cs.LG2026
H2LooP Spark Preview: Continual Pretraining of Large Language Models for Low-Level Embedded Systems Code
Amit Singh, Vedant Nipane, Pulkit Agrawal +2
Large language models (LLMs) demonstrate strong code generation abilities in general-purpose programming languages but remain limited in specialized domains such as low-level embed…
cs.CV2024
Fluid Dynamic DNNs for Reliable and Adaptive Distributed Inference on Edge Devices
Lei Xun, Mingyu Hu, Hengrui Zhao +3
Distributed inference is a popular approach for efficient DNN inference at the edge. However, traditional Static and Dynamic DNNs are not distribution-friendly, causing system reli…