Showing cs.AIShow all
2 papers · 1 filter
cs.AI2025
Boosting Universal LLM Reward Design through Heuristic Reward Observation Space Evolution
Zen Kit Heng, Zimeng Zhao, Tianhao Wu +4
Large Language Models (LLMs) are emerging as promising tools for automated reinforcement learning (RL) reward design, owing to their robust capabilities in commonsense reasoning an…
cs.AI2024
Fast Peer Adaptation with Context-aware Exploration
Long Ma, Yuanfei Wang, Fangwei Zhong +2
Fast adapting to unknown peers (partners or opponents) with different strategies is a key challenge in multi-agent games. To do so, it is crucial for the agent to probe and identif…