3 papers
cs.LG2026
DiLaServe: High SLO Attainment Serving for Diffusion Language Models
Tzu-Tao Chang, Benjamin Yuanyang Hong, Kiet Pham +1
Diffusion language models (DLMs) have recently emerged as a promising alternative to conventional autoregressive language models. By generating multiple tokens in parallel during e…
cs.CV2025
LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models
Tzu-Tao Chang, Shivaram Venkataraman
Cross-attention is commonly adopted in multimodal large language models (MLLMs) for integrating visual information into the language backbone. However, in applications with large v…
cs.DC2025
Eva: Cost-Efficient Cloud-Based Cluster Scheduling
Tzu-Tao Chang, Shivaram Venkataraman
Cloud computing offers flexibility in resource provisioning, allowing an organization to host its batch processing workloads cost-efficiently by dynamically scaling the size and co…