Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Beyond the Stability-Exploration Dilemma: Environmental Regularization for LLM Policy Optimization
Xianlei Zhou, Xiangdi Meng, Yu He +7
Policy optimization (PO) for Large Language Models faces a stability--exploration trade-off, currently mediated by an action-side Policy-KL regularizer. This puts practitioners in…
cs.CL2025
MARS: Multi-Agent Adaptive Reasoning with Socratic Guidance for Automated Prompt Optimization
Jian Zhang, Zhangqi Wang, Haiping Zhu +6
Large language models (LLMs) typically operate in a question-answering paradigm, where the quality of the input prompt critically affects the response. Automated Prompt Optimizatio…