5 papers
AutoSynth: Automated Workflow Optimization for High-Quality Synthetic Dataset Generation via Monte Carlo Tree Search
Shuzhen Bi, Chang Song, Siyu Song +5
Supervised fine-tuning (SFT) of large language models (LLMs) for specialized tasks requires high-quality datasets, but manual curation is prohibitively expensive. Synthetic data ge…
Lyte Quorum: Off-Chain Ready Smart Contract Hosted with Choice
Hao Hao, Dahlia Malkhi, Maofan Yin +1
This paper introduces Lyquor, a decentralized platform that reimagines blockchain infrastructure through a service-centric model where nodes selectively host smart contracts (calle…
SID: Benchmarking Guided Instruction Capabilities in STEM Education with a Socratic Interdisciplinary Dialogues Dataset
Mei Jiang, Houping Yue, Bingdong Li +4
Fostering students' abilities for knowledge integration and transfer in complex problem-solving scenarios is a core objective of modern education, and interdisciplinary STEM is a k…
ELMES: An Automated Framework for Evaluating Large Language Models in Educational Scenarios
Shou'ang Wei, Xinyun Wang, Shuzhen Bi +9
The emergence of Large Language Models (LLMs) presents transformative opportunities for education, generating numerous novel application scenarios. However, significant challenges…
Cultivating Helpful, Personalized, and Creative AI Tutors: A Framework for Pedagogical Alignment using Reinforcement Learning
Siyu Song, Wentao Liu, Ye Lu +8
The integration of large language models (LLMs) into education presents unprecedented opportunities for scalable personalized learning. However, standard LLMs often function as gen…