1 citations · 1 across the 8 of their papers we have counts for
7 papers
KAIROS: Stateful, Context-Aware Power-Efficient Agentic Inference Serving
Yichao Yuan, Mosharaf Chowdhury, Nishil Talati
Power has become a central bottleneck for AI inference. This problem is becoming more urgent as agentic AI emerges as a major workload class, yet prior power-management techniques…
BlazingAML: High-Throughput Anti-Money Laundering (AML) via Multi-Stage Graph Mining
Haojie Ye, Arjun Laxman, Yichao Yuan +2
Money laundering detection faces challenges due to excessive false positives and inadequate adaptation to sophisticated multi-stage schemes that exploit modern financial networks.…
ZKProphet: Understanding Performance of Zero-Knowledge Proofs on GPUs
Tarunesh Verma, Yichao Yuan, Nishil Talati +1
Zero-Knowledge Proofs (ZKP) are protocols which construct cryptographic proofs to demonstrate knowledge of a secret input in a computation without revealing any information about t…
MoE-Lens: Towards the Hardware Limit of High-Throughput MoE LLM Serving Under Resource Constraints
Yichao Yuan, Lin Ma, Nishil Talati
Mixture of Experts (MoE) LLMs, characterized by their sparse activation patterns, offer a promising approach to scaling language models while avoiding proportionally increasing the…
MoDM: Efficient Serving for Image Generation via Mixture-of-Diffusion Models
Yuchen Xia, Divyam Sharma, Yichao Yuan +2
Diffusion-based text-to-image generation models trade latency for quality: small models are fast but generate lower-quality images, while large models produce better images but are…
Vortex: Overcoming Memory Capacity Limitations in GPU-Accelerated Large-Scale Data Analytics
Yichao Yuan, Advait Iyer, Lin Ma +1
Despite the high computational throughput of GPUs, limited memory capacity and bandwidth-limited CPU-GPU communication via PCIe links remain significant bottlenecks for acceleratin…