6 papers
Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective
Feng Zhang, Xinhong Ma, Ziqiang Dong +5
Group Relative Policy Optimization (GRPO) is one of the most widely adopted RLVR algorithms for post-training large language models on reasoning tasks. We first show that GRPO admi…
Metis: Learning to Jailbreak LLMs via Self-Evolving Metacognitive Policy Optimization
Huilin Zhou, Jian Zhao, Yilu Zhong +7
Red teaming is critical for uncovering vulnerabilities in Large Language Models (LLMs). While automated methods have improved scalability, existing approaches often rely on static…
2K-Characters-10K-Stories: A Quality-Gated Stylized Narrative Dataset with Disentangled Control and Sequence Consistency
Xingxi Yin, Yicheng Li, Gong Yan +5
Sequential identity consistency under precise transient attribute control remains a long-standing challenge in controllable visual storytelling. Existing datasets lack sufficient f…
EvoCurr: Self-evolving Curriculum with Behavior Code Generation for Complex Decision-making
Yang Cheng, Zilai Wang, Weiyu Ma +3
Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse domains, including programming, planning, and decision-making. However, their performance ofte…
SMAC-R1: The Emergence of Intelligence in Decision-Making Tasks
Yue Deng, Weiyu Ma, Yuxin Fan +4
StarCraft Multi-Agent Challenge (SMAC) has been one of the most commonly used experimental environments in multi-agent reinforcement learning (MARL), where the specific task is to…
SMAC-Hard: Enabling Mixed Opponent Strategy Script and Self-play on SMAC
Yue Deng, Yan Yu, Weiyu Ma +4
The availability of challenging simulation environments is pivotal for advancing the field of Multi-Agent Reinforcement Learning (MARL). In cooperative MARL settings, the StarCraft…