collaborators

5 papers

cs.AI2026

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning

Zibin Meng, Zhenyu Zhao, Chunqiang Run

Reinforcement learning with verifiable rewards is a powerful paradigm for eliciting reasoning in large language models, yet it suffers from severe reward sparsity on competition-le…

cs.LG2026

CRAFT: Counterfactual Credit Assignment from Free Sibling Rollouts for Self-Distilled Agentic Reinforcement Learning

Zibin Meng, Kani Chen

Self-distilled agentic reinforcement learning augments trajectory-level reward with a token-level distillation loss, using as its teacher the same policy conditioned on privileged…

cs.LG2026

Depth-Entropy Guided Sampling for Training-Free LLM Reasoning

Zibin Meng, Peng Xie, Kani Chen

Reinforcement learning (RL) has become the dominant paradigm for improving the reasoning capabilities of large language models, but it requires expensive training, curated data, an…

cs.AI2026

PsyAgent: Constructing Human-like Agents Based on Psychological Modeling and Contextual Interaction

Zibin Meng, Kani Chen

Human-like agents must express stable dispositions while adapting to roles, relationships, and norms. We present PsyAgent, a schema-first framework that operationalizes the trait-c…

cs.CV2024

SocialGPT: Prompting LLMs for Social Relation Reasoning via Greedy Segment Optimization

Wanhua Li, Zibin Meng, Jiawei Zhou +3

Social relation reasoning aims to identify relation categories such as friends, spouses, and colleagues from images. While current methods adopt the paradigm of training a dedicate…