2 papers
cs.LG2026
MASPOB: Bandit-Based Prompt Optimization for Multi-Agent Systems with Graph Neural Networks
Zhi Hong, Qian Zhang, Jiahang Sun +5
Large Language Models (LLMs) have achieved great success in many real-world applications, especially the one serving as the cognitive backbone of Multi-Agent Systems (MAS) to orche…
cs.LG2026
Linear and Neural Dueling Bandits with Delayed Feedback
Xiangyi Wang, Pingchen Lu, Jie Mao +4
Contextual dueling bandits form a cornerstone of preference-based decision-making, with critical applications in recommender systems and large language model alignment. However, st…