2.1k citations · 2.1k across the 31 of their papers we have counts for
4 papers · 2 filters
SAVOIR: Learning Social Savoir-Faire via Shapley-based Reward Attribution
Xiachong Feng, Yi Jiang, Xiaocheng Feng +9
Social intelligence, the ability to navigate complex interpersonal interactions, presents a fundamental challenge for language agents. Training such agents via reinforcement learni…
Stratagem: Learning Transferable Reasoning via Trajectory-Modulated Game Self-Play
Xiachong Feng, Deyi Yin, Xiaocheng Feng +9
Games offer a compelling paradigm for developing general reasoning capabilities in language models, as they naturally demand strategic planning, probabilistic inference, and adapti…
Not All Tokens See Equally: Perception-Grounded Policy Optimization for Large Vision-Language Models
Zekai Ye, Qiming Li, Xiaocheng Feng +6
While Reinforcement Learning from Verifiable Rewards (RLVR) has advanced reasoning in Large Vision-Language Models (LVLMs), prevailing frameworks suffer from a foundational methodo…
PERSONA: Dynamic and Compositional Inference-Time Personality Control via Activation Vector Algebra
Xiachong Feng, Liang Zhao, Weihong Zhong +5
Current methods for personality control in Large Language Models rely on static prompting or expensive fine-tuning, failing to capture the dynamic and compositional nature of human…