Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE
Haozhan Tang, Zerui Wang, Yuxian Gu +2
Modern LLMs are increasingly deployed in long-context applications such as retrieval-augmented generation, repository-level coding, and agentic workflows whose accumulated reasonin…
cs.LG2025
Semantic-Aware Scheduling for GPU Clusters with Large Language Models
Zerui Wang, Qinghao Hu, Ana Klimovic +4
Deep learning (DL) schedulers are pivotal in optimizing resource allocation in GPU clusters, but operate with a critical limitation: they are largely blind to the semantic context…
cs.LG2024
PackMamba: Efficient Processing of Variable-Length Sequences in Mamba training
Haoran Xu, Ziqian Liu, Rong Fu +5
With the evolution of large language models, traditional Transformer models become computationally demanding for lengthy sequences due to the quadratic growth in computation with r…