5 papers
Host-Side Telemetry for Performance Diagnosis in Cloud and HPC GPU Infrastructure
Erfan Darzi, Aldo Pareja, Shreeanant Bharadwaj
Diagnosing GPU tail latency spikes in cloud and HPC infrastructure is critical for maintaining performance predictability and resource utilization, yet existing monitoring tools la…
Predictable LLM Serving on GPU Clusters
Erfan Darzi, Shreeanant Bharadwaj, Sree Bhargavi Balija
Latency-sensitive inference on shared A100 clusters often suffers noisy-neighbor interference on the PCIe fabric, inflating tail latency and SLO violations. We present a fabric-agn…
The Trust Fabric: Decentralized Interoperability and Economic Coordination for the Agentic Web
Sree Bhargavi Balija, Rekha Singal, Ramesh Raskar +4
The fragmentation of AI agent ecosystems has created urgent demands for interoperability, trust, and economic coordination that current protocols -- including MCP (Hou et al., 2025…
Foundation AI Model for Medical Image Segmentation
Rina Bao, Erfan Darzi, Sheng He +6
Foundation models refer to artificial intelligence (AI) models that are trained on massive amounts of data and demonstrate broad generalizability across various tasks with high acc…
A Survey for Large Language Models in Biomedicine
Chong Wang, Mengyao Li, Junjun He +14
Recent breakthroughs in large language models (LLMs) offer unprecedented natural language understanding and generation capabilities. However, existing surveys on LLMs in biomedicin…