11 papers
Themis: Software-Defined Hardware Prefetching
Keisuke Kamahori, Neil Adit, Kan Zhu +13
Data cache misses represent a significant portion of stall cycles in datacenter workloads. Hardware prefetchers that reduce such stalls by fetching data ahead of time have become i…
M*: A Modular, Extensible, Serving System for Multimodal Models
Atindra Jha, Naomi Sagan, Keisuke Kamahori +9
We are entering a new era of composite model architectures that integrate diverse components such as vision encoders, language backbones, diffusion and flow heads, audio codecs, ac…
MURMUR: An Efficient Inference System for Long-Form ASR
Wei-Tzu Lee, Keisuke Kamahori, Baris Kasikci
Long-form automatic speech recognition (ASR) requires both high accuracy and low latency, but existing systems force a trade-off between the two. Chunk-based pipelines process audi…
TeleRAG: Efficient Retrieval-Augmented Generation Inference with Lookahead Retrieval
Chien-Yu Lin, Keisuke Kamahori, Yiyu Liu +11
Retrieval-augmented generation (RAG) extends large language models (LLMs) with external data sources to enhance factual correctness and domain coverage. Modern RAG pipelines rely o…
VibeServe: Can AI Agents Build Bespoke LLM Serving Systems?
Keisuke Kamahori, Shihang Li, Simon Peter +1
For years, we have built LLM serving systems like any other critical infrastructure: a single general-purpose stack, hand-tuned over many engineer-years, meant to support every mod…
VoxServe: Streaming-Centric Serving System for Speech Language Models
Keisuke Kamahori, Wei-Tzu Lee, Atindra Jha +4
Deploying modern Speech Language Models (SpeechLMs) in streaming settings requires systems that provide low latency, high throughput, and strong guarantees of streamability. Existi…