2 citations · 3 across the 6 of their papers we have counts for
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2025
Trainable Dynamic Mask Sparse Attention
Jingze Shi, Yifan Wu, Yiran Peng +4
The increasing demand for long-context modeling in large language models (LLMs) is bottlenecked by the quadratic complexity of the standard self-attention mechanism. The community…
cs.AI2025
Concise Reasoning, Big Gains: Pruning Long Reasoning Trace with Difficulty-Aware Prompting
Yifan Wu, Jingze Shi, Bingheng Wu +4
Existing chain-of-thought (CoT) distillation methods can effectively transfer reasoning abilities to base models but suffer from two major limitations: excessive verbosity of reaso…