2 papers
cs.LG2026
M*: A Modular, Extensible, Serving System for Multimodal Models
Atindra Jha, Naomi Sagan, Keisuke Kamahori +9
We are entering a new era of composite model architectures that integrate diverse components such as vision encoders, language backbones, diffusion and flow heads, audio codecs, ac…
cs.LG2026
VoxServe: Streaming-Centric Serving System for Speech Language Models
Keisuke Kamahori, Wei-Tzu Lee, Atindra Jha +4
Deploying modern Speech Language Models (SpeechLMs) in streaming settings requires systems that provide low latency, high throughput, and strong guarantees of streamability. Existi…