activity
20242026
most citedStreamingVLM: Real-Time Understanding for Infinite Video Streams

1 citations · 1 across the 3 of their papers we have counts for

collaborators

9 papers

cs.CL2026

SparDA: Sparse Decoupled Attention for Efficient Long-Context LLM Inference

Yaosheng Fu, Guangxuan Xiao, Xin Dong +2

Sparse attention reduces compute and memory bandwidth for long-context LLM inference. However, two key challenges remain: (1) KV cache capacity still grows with sequence length, an…

cs.CV20261 cited

StreamingVLM: Real-Time Understanding for Infinite Video Streams

Ruyi Xu, Guangxuan Xiao, Yukang Chen +3

Vision-language models (VLMs) could power real-time assistants and autonomous agents, but they face a critical challenge: understanding near-infinite video streams without escalati…

cs.CL2026

BLASST: Dynamic BLocked Attention Sparsity via Softmax Thresholding

Jiayi Yuan, Cameron Shinn, Kai Xu +19

The growing demand for long-context inference capabilities in Large Language Models (LLMs) has intensified the computational and memory bottlenecks inherent to the self-attention m…

cs.LG2025

Optimizing Mixture of Block Attention

Guangxuan Xiao, Junxian Guo, Kasra Mazaheri +1

Mixture of Block Attention (MoBA) (Lu et al., 2025) is a promising building block for efficiently processing long contexts in LLMs by enabling queries to sparsely attend to a small…

cs.LG2025

Twilight: Adaptive Attention Sparsity with Hierarchical Top- Pruning

Chaofan Lin, Jiaming Tang, Shuo Yang +6

Leveraging attention sparsity to accelerate long-context large language models (LLMs) has been a hot research topic. However, current algorithms such as sparse attention or key-val…

cs.CL2025

LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention

Shang Yang, Junxian Guo, Haotian Tang +7

Large language models (LLMs) have shown remarkable potential in processing long sequences and complex reasoning tasks, yet efficiently serving these models remains challenging due…