7 papers
Freshness-Aware Prioritized Experience Replay for LLM/VLM Reinforcement Learning
Weiyu Ma, Yongcheng Zeng, Yan Song +4
Reinforcement Learning (RL) has achieved impressive success in post-training Large Language Models (LLMs) and Vision-Language Models (VLMs), with on-policy algorithms such as PPO,…
Adaptive Command: Real-Time Policy Adjustment via Language Models in StarCraft II
Weiyu Ma, Dongyu Xu, Shu Lin +2
We present Adaptive Command, a novel framework integrating large language models (LLMs) with behavior trees for real-time strategic decision-making in StarCraft II. Our system focu…
Evolving LLMs' Self-Refinement Capability via Synergistic Training-Inference Optimization
Yongcheng Zeng, Xinyu Cui, Xuanfa Jin +11
Self-Refinement refers to a model's ability to revise its own responses to produce improved outputs. This capability can also serve as a fundamental mechanism for Self-Improvement,…
EvoCurr: Self-evolving Curriculum with Behavior Code Generation for Complex Decision-making
Yang Cheng, Zilai Wang, Weiyu Ma +3
Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse domains, including programming, planning, and decision-making. However, their performance ofte…
TacticCraft: Natural Language-Driven Tactical Adaptation for StarCraft II
Weiyu Ma, Jiwen Jiang, Haobo Fu +1
We present an adapter-based approach for tactical conditioning of StarCraft II AI agents. Current agents, while powerful, lack the ability to adapt their strategies based on high-l…
SMAC-R1: The Emergence of Intelligence in Decision-Making Tasks
Yue Deng, Weiyu Ma, Yuxin Fan +4
StarCraft Multi-Agent Challenge (SMAC) has been one of the most commonly used experimental environments in multi-agent reinforcement learning (MARL), where the specific task is to…