collaborators

12 papers

cs.CV2026

FG-Attn: Leveraging Fine-Grained Sparse Attention in Video Diffusion Models

Sankeerth Durvasula, Kavya Sreedhar, Zain Moustafa +6

Using diffusion transformers for media generation may require evaluating attention over extremely long sequences, with attention layers accounting for the majority of generation la…

cs.LG2026

How Far Can Disaggregation Go? A Design-Space Exploration of Attention-FFN Disaggregation for Efficient MoE LLM Serving

Hanjiang Wu, Abhimanyu Rajeshkumar Bambhaniya, Sarbartha Banerjee +9

Modern large language model (LLM) inference has progressively disaggregated to keep pace with growing model sizes and tight TTFT and TPOT service-level objectives: from chunked-pre…

cs.AI2026

JAXBench: Benchmarking Autonomous TPU Kernel Optimization

Arya Tschand, Charles Hong, Julian Walker +7

Rigorous benchmarks have driven progress in autonomous GPU kernel performance optimization by establishing a shared target to hillclimb on, but no equivalent exists for TPUs. We pr…

cs.AR2026

MIST: A Co-Design Framework for Heterogeneous, Multi-Stage LLM Inference

Abhimanyu Rajeshkumar Bambhaniya, Hanjiang Wu, Suvinay Subramanian +8

Modern LLM serving now spans multi-stage pipelines including RAG retrieval and KV cache reuse, each with distinct compute, memory, and latency demands. Inference engines expose a l…

cs.AI2026

Planned Diffusion

Daniel Israel, Tian Jin, Ellie Cheng +4

Most large language models are autoregressive: they generate tokens one at a time. Discrete diffusion language models can generate multiple tokens in parallel, but sampling from th…

cs.PF2026

Characterizing VLA Models: Identifying the Action Generation Bottleneck for Edge AI Architectures

Manoj Vishwanathan, Suvinay Subramanian, Anand Raghunathan

Vision-Language-Action (VLA) models are an emerging class of workloads critical for robotics and embodied AI at the edge. As these models scale, they demonstrate significant capabi…