2 papers
cs.LG2026
Large Language Model Prompt Datasets: An In-depth Analysis and Insights
Yuanming Zhang, Yan Lin, Arijit Khan +1
We compile 129 heterogeneous LLM prompt datasets (>1.22 TB, >673M instances) into a structured taxonomy and conduct a multi-level linguistic analysis (lexical, syntactic, and seman…
cs.CV2025
EduPersona: Benchmarking Subjective Ability Boundaries of Virtual Student Agents
Buyuan Zhu, Shiyu Hu, Yiping Ma +2
As large language models are increasingly integrated into education, virtual student agents are becoming vital for classroom simulation and teacher training. Yet their classroom-or…