2 papers
cs.CL2026
IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse
Yushi Bai, Qian Dong, Ting Jiang +5
Long-context agentic workflows have emerged as a defining use case for large language models, making attention efficiency critical for both inference speed and serving cost. Sparse…
cs.LG2025
SIRI: Scaling Iterative Reinforcement Learning with Interleaved Compression
Haoming Wen, Yushi Bai, Juanzi Li +1
We introduce SIRI, Scaling Iterative Reinforcement Learning with Interleaved Compression, a simple yet effective RL approach for Large Reasoning Models (LRMs) that enables more eff…