2 citations · 5 across the 32 of their papers we have counts for
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
SRPO: Self-Reflective Policy Optimization for Long-Horizon Reasoning
Jialong Liu, Yuling Shi, Ning Yang +2
Self-reflection is a powerful mechanism for credit assignment in human learning, converting sparse outcome feedback into actionable guidance. However, its potential for post-traini…
cs.AI2026
Reward-Decomposed Reinforcement Learning for Immersive Video Role-Playing
Miao Wang, Yuling Shi, Yijiang Li +8
Text-based role-playing models can imitate character styles, but often fail to capture scene atmosphere and evolving tension, which are crucial for immersive applications such as V…