activity
20232025
most citedMAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning

21 citations · 28 across the 7 of their papers we have counts for

collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2025

COIG-P: A High-Quality and Large-Scale Chinese Preference Dataset for Alignment with Human Values

P Team, Siwei Wu, Jincheng Ren +29

Aligning large language models (LLMs) with human preferences has achieved remarkable success. However, existing Chinese preference datasets are limited by small scale, narrow domai…

cs.CL2024

Overview of the NLPCC 2024 Shared Task on Chinese Metaphor Generation

Xingwei Qu, Ge Zhang, Siwei Wu +2

This paper presents the results of the shared task on Chinese metaphor generation, hosted at the 13th CCF Conference on Natural Language Processing and Chinese Computing (NLPCC 202…

cs.CL20243 cited

D-CPT Law: Domain-specific Continual Pre-Training Scaling Law for Large Language Models

Haoran Que, Jiaheng Liu, Ge Zhang +13

Continual Pre-Training (CPT) on Large Language Models (LLMs) has been widely used to expand the model's fundamental understanding of specific downstream domains (e.g., math and cod…

cs.CL20243 cited

CMDAG: A Chinese Metaphor Dataset with Annotated Grounds as CoT for Boosting Metaphor Generation

Yujie Shao, Xinrong Yao, Xingwei Qu +5

Metaphor is a prominent linguistic device in human language and literature, as they add color, imagery, and emphasis to enhance effective communication. This paper introduces a lar…

cs.CL202321 cited

MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning

Xiang Yue, Xingwei Qu, Ge Zhang +5

We introduce MAmmoTH, a series of open-source large language models (LLMs) specifically tailored for general math problem-solving. The MAmmoTH models are trained on MathInstruct, o…