collaborators

6 papers

cs.LG2026

GeoMin: Data-Efficient Semi-Supervised RLVR via Geometric Distribution Modeling

Guangcheng Zhu, Shenzhi Yang, Haobo Wang +9

Reinforcement learning with verifiable rewards (RLVR) significantly advances LLM reasoning, yet it faces a dilemma: standard supervised scaling is throttled by high annotation cost…

cs.LG2026

Smart Picks in the Dark: Towards Efficient RLVR for Reasoning via Tracing Metacognitive Pivots

Guangcheng Zhu, Shenzhi Yang, Haobo Wang +7

Reinforcement learning with verifiable rewards (RLVR) has greatly advanced large reasoning models (LRMs), but it requires timely training on a huge fully-annotated dataset. To this…

cs.GT2026

Learning to cooperate with emergent reputation via multi-agent reinforcement learning

Xinwei Song, Yizhe Huang, Dengji Zhao +1

Reputation, the aggregation of peer assessments diffused through social networks, is a pivotal mechanism for promoting cooperation in social dilemmas ubiquitous to distributed mult…

cs.AI2025

ValuePilot: A Two-Phase Framework for Value-Driven Decision-Making

Yitong Luo, Ziang Chen, Hou Hei Lam +4

Personalized decision-making is essential for human-AI interaction, enabling AI agents to act in alignment with individual users' value preferences. As AI systems expand into real-…

cs.AI2025

ToMPO: Training LLM Strategic Decision Making from a Multi-Agent Perspective

Yiwen Zhang, Ziang Chen, Fanqi Kong +2

Large Language Models (LLMs) have been used to make decisions in complex scenarios, where they need models to think deeply, reason logically, and decide wisely. Many existing studi…

cs.AI2025

ValuePilot: A Two-Phase Framework for Value-Driven Decision-Making

Yitong Luo, Hou Hei Lam, Ziang Chen +2

Despite recent advances in artificial intelligence (AI), it poses challenges to ensure personalized decision-making in tasks that are not considered in training datasets. To addres…