evaluation benchmark 1large language models 1multi-turn dialogue 1role-playing agents 1user simulation 1
From the 1 of 7 linked papers with an AI index.
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Adaptive Supervised Anchoring for On-Policy Self-Distillation
Meilin Yang, Zixuan Ding, Jianhao Nie +5
On-policy self-distillation (OPSD) adapts a language model by distilling guidance from a frozen teacher on trajectories sampled from the student. Its effectiveness, however, depend…
cs.LG2026
LABO: LLM-Accelerated Bayesian Optimization through Broad Exploration and Selective Experimentation
Zhuo Chen, Xinzhe Yuan, Jianshu Zhang +8
The high cost and data scarcity in scientific exploration have motivated the use of large language models (LLMs) as knowledge-driven components in Bayesian optimization (BO). Howev…