Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Behavioural Signatures of Risk-Sensitive Decision-Making in Large Language Models
Xuankun Rong, Wenke Huang, Bo Du +2
As large language models (LLMs) are increasingly used in decision support, it is important to understand whether their choices under uncertainty exhibit stable and interpretable be…
cs.AI2025
MAPO: Mixed Advantage Policy Optimization
Wenke Huang, Quan Zhang, Yiyang Fang +11
Recent advances in reinforcement learning for foundation models, such as Group Relative Policy Optimization (GRPO), have significantly improved the performance of foundation models…