34 citations · 38 across the 11 of their papers we have counts for
4 papers · 1 filter
Mesh-Attention: A New Communication-Efficient Distributed Attention with Improved Data Locality
Sirui Chen, Jingji Chen, Siqi Zhu +3
Distributed attention is essential for scaling large language models (LLMs) to long contexts, yet existing methods either have limited parallelism or incur high communication costs…
Arrows of Math Reasoning Data Synthesis for Large Language Models: Diversity, Complexity and Correctness
Sirui Chen, Changxin Tian, Binbin Hu +4
Enhancing the mathematical reasoning of large language models (LLMs) demands high-quality training data, yet conventional methods face critical challenges in scalability, cost, and…
OpenGT: A Comprehensive Benchmark For Graph Transformers
Jiachen Tang, Zhonghao Wang, Sirui Chen +3
Graph Transformers (GTs) have recently demonstrated remarkable performance across diverse domains. By leveraging attention mechanisms, GTs are capable of modeling long-range depend…
Rankformer: A Graph Transformer for Recommendation based on Ranking Objective
Sirui Chen, Shen Han, Jiawei Chen +6
Recommender Systems (RS) aim to generate personalized ranked lists for each user and are evaluated using ranking metrics. Although personalized ranking is a fundamental aspect of R…