7 citations · 8 across the 2 of their papers we have counts for
3 papers
cs.LG2024★ 1 cited
MindStar: Enhancing Math Reasoning in Pre-trained LLMs at Inference Time
Jikun Kang, Xin Zhe Li, Xi Chen +9
Although Large Language Models (LLMs) achieve remarkable performance across various tasks, they often struggle with complex reasoning tasks, such as answering mathematical question…
cs.CL2023★ 7 cited
PanGu-Σ: Towards Trillion Parameter Language Model with Sparse Heterogeneous Computing
Xiaozhe Ren, Pingyi Zhou, Xinfan Meng +14
The scaling of large language models has greatly improved natural language understanding, generation, and reasoning. In this work, we develop a system that trained a trillion-param…
cs.LG2023
ArCL: Enhancing Contrastive Learning with Augmentation-Robust Representations
Xuyang Zhao, Tianqi Du, Yisen Wang +2
Self-Supervised Learning (SSL) is a paradigm that leverages unlabeled data for model training. Empirical studies show that SSL can achieve promising performance in distribution shi…