large language models 2data synthesis 1humanities 1humanities and social sciences 1instruction tuning 1preference alignment 1quality evaluation 1social sciences 1synthetic data generation 1
From the 2 of 3 linked papers with an AI index.
3 papers
cs.CL2026
HSS-Synth: Humanities and Social Sciences Data Synthesis for LLMs
Ru Peng, Tianyu Zhao, Xijun Gu +9
The paper introduces HSS-Synth, a pipeline that creates high‑quality instruction‑tuning data for large language models in the humanities and social sciences by generating seed docu…
cs.CL2026
BridgeAlign: Bridging Preference Alignment for Humanities and Social Sciences
Ru Peng, Haokai Xu, Xijun Gu +11
BridgeAlign introduces a three-stage pipeline that creates and uses synthetic preference data to align large language models with nuanced quality judgments in humanities and social…
cs.CL2025
DataMan: Data Manager for Pre-training Large Language Models
Ru Peng, Kexin Yang, Yawen Zeng +3
The performance emergence of large language models (LLMs) driven by data scaling laws makes the selection of pre-training data increasingly important. However, existing methods rel…