6 papers
GeoMin: Data-Efficient Semi-Supervised RLVR via Geometric Distribution Modeling
Guangcheng Zhu, Shenzhi Yang, Haobo Wang +9
Reinforcement learning with verifiable rewards (RLVR) significantly advances LLM reasoning, yet it faces a dilemma: standard supervised scaling is throttled by high annotation cost…
Smart Picks in the Dark: Towards Efficient RLVR for Reasoning via Tracing Metacognitive Pivots
Guangcheng Zhu, Shenzhi Yang, Haobo Wang +7
Reinforcement learning with verifiable rewards (RLVR) has greatly advanced large reasoning models (LRMs), but it requires timely training on a huge fully-annotated dataset. To this…
Learning to cooperate with emergent reputation via multi-agent reinforcement learning
Xinwei Song, Yizhe Huang, Dengji Zhao +1
Reputation, the aggregation of peer assessments diffused through social networks, is a pivotal mechanism for promoting cooperation in social dilemmas ubiquitous to distributed mult…
ValuePilot: A Two-Phase Framework for Value-Driven Decision-Making
Yitong Luo, Ziang Chen, Hou Hei Lam +4
Personalized decision-making is essential for human-AI interaction, enabling AI agents to act in alignment with individual users' value preferences. As AI systems expand into real-…
ToMPO: Training LLM Strategic Decision Making from a Multi-Agent Perspective
Yiwen Zhang, Ziang Chen, Fanqi Kong +2
Large Language Models (LLMs) have been used to make decisions in complex scenarios, where they need models to think deeply, reason logically, and decide wisely. Many existing studi…
ValuePilot: A Two-Phase Framework for Value-Driven Decision-Making
Yitong Luo, Hou Hei Lam, Ziang Chen +2
Despite recent advances in artificial intelligence (AI), it poses challenges to ensure personalized decision-making in tasks that are not considered in training datasets. To addres…