12 papers
Elmes*: Automated Construction of Fine-Grained Evaluation Rubrics for Large Language Models in Long-Tail Educational Scenarios
Tao Liu, Ye Lu, Ruohua Zhang +4
Evaluating large language models (LLMs) for education requires measuring how models teach, not only what they know. Existing benchmarks emphasize domain-general correctness or depe…
Relation Reasoning with LLMs in Expensive Optimization
Ye Lu, Bingdong Li, Aimin Zhou +1
Expensive optimization problems (EOPs) are black-box tasks with costly objective evaluations and no gradient access, making the evaluation budget the key bottleneck. Surrogate-assi…
Automating Skill Acquisition through Large-Scale Mining of Open-Source Agentic Repositories: A Framework for Multi-Agent Procedural Knowledge Extraction
Shuzhen Bi, Mengsong Wu, Hao Hao +5
The transition from monolithic large language models (LLMs) to modular, skill-equipped agents represents a fundamental architectural shift in artificial intelligence deployment. Wh…
Scaling Laws for Educational AI Agents
Mengsong Wu, Hao Hao, Shuzhen Bi +5
While scaling laws for Large Language Models (LLMs) have been extensively studied along dimensions of model parameters, training data, and compute, the scaling behavior of LLM-base…
See and Remember: A Multimodal Agent for Web Traversal
Xinjun Wang, Shengyao Wang, Aimin Zhou +1
Autonomous web navigation requires agents to perceive complex visual environments and maintain long-term context, yet current Large Language Model (LLM) based agents often struggle…
AutoSynth: Automated Workflow Optimization for High-Quality Synthetic Dataset Generation via Monte Carlo Tree Search
Shuzhen Bi, Chang Song, Siyu Song +5
Supervised fine-tuning (SFT) of large language models (LLMs) for specialized tasks requires high-quality datasets, but manual curation is prohibitively expensive. Synthetic data ge…