2 papers
cs.LG2026
Sampling for Quality: Training-Free Reward-Guided LLM Decoding via Sequential Monte Carlo
Jelena Markovic-Voronov, Wenhui Zhu, Bo Long +5
We introduce a principled probabilistic framework for reward-guided decoding in large language models, addressing the limitations of standard decoding methods that optimize token-l…
cs.AI2024
SMAC-Hard: Enabling Mixed Opponent Strategy Script and Self-play on SMAC
Yue Deng, Yan Yu, Weiyu Ma +4
The availability of challenging simulation environments is pivotal for advancing the field of Multi-Agent Reinforcement Learning (MARL). In cooperative MARL settings, the StarCraft…