activity
20242026
most citedNMRTrans: Structure Elucidation from Experimental NMR Spectra via Set Transformers

1 citations · 1 across the 3 of their papers we have counts for

collaborators
Showing cs.CLShow all

6 papers · 1 filter

cs.CL2025

Topic Over Source: The Key to Effective Data Mixing for Language Models Pre-training

Jiahui Peng, Xinlin Zhuang, Jiantao Qiu +4

The performance of large language models (LLMs) is significantly affected by the quality and composition of their pre-training data, which is inherently diverse, spanning various l…

cs.CL2025

Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models

Xinlin Zhuang, Jiahui Peng, Ren Ma +7

The composition of pre-training datasets for large language models (LLMs) remains largely undisclosed, hindering transparency and efforts to optimize data quality, a critical drive…

cs.CL2025

Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration

Tianyi Bai, Ling Yang, Zhen Hao Wong +9

Efficient data selection is crucial to accelerate the pretraining of language model (LMs). While various methods have been proposed to enhance data efficiency, limited research has…

cs.CL2025

Not All Documents Are What You Need for Extracting Instruction Tuning Data

Chi Zhang, Huaping Zhong, Hongtao Li +11

Instruction tuning improves the performance of large language models (LLMs), but it heavily relies on high-quality training data. Recently, LLMs have been used to synthesize instru…

cs.CL2025

BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models

Xu Huang, Wenhao Zhu, Hanxu Hu +4

Previous multilingual benchmarks focus primarily on simple understanding tasks, but for large language models(LLMs), we emphasize proficiency in instruction following, reasoning, l…

cs.CL2025

Large Language Models Meet Symbolic Provers for Logical Reasoning Evaluation

Chengwen Qi, Ren Ma, Bowen Li +5

First-order logic (FOL) reasoning, which involves sequential deduction, is pivotal for intelligent systems and serves as a valuable task for evaluating reasoning capabilities, part…