3 papers
cs.AI2026
COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization
Pingchen Lu, Xiangyi Wang, Xiang Li +6
Large language model (LLM) agents can benefit from reusable skills distilled from prior task experience, yet existing skill optimization methods often rely on costly execution-base…
cs.LG2026
Linear and Neural Dueling Bandits with Delayed Feedback
Xiangyi Wang, Pingchen Lu, Jie Mao +4
Contextual dueling bandits form a cornerstone of preference-based decision-making, with critical applications in recommender systems and large language model alignment. However, st…
cs.LG2026
MASPOB: Bandit-Based Prompt Optimization for Multi-Agent Systems with Graph Neural Networks
Zhi Hong, Qian Zhang, Jiahang Sun +5
Large Language Models (LLMs) have achieved great success in many real-world applications, especially the one serving as the cognitive backbone of Multi-Agent Systems (MAS) to orche…