2 papers
cs.LG2026
The Two-Stage Decision-Sampling Hypothesis: Understanding the Emergence of Self-Reflection in RL-Trained LLMs
Zibo Zhao, Yuanting Zha, Haipeng Zhang +1
Self-reflection capabilities emerge in Large Language Models after RL post-training, with multi-turn RL achieving substantial gains over SFT counterparts. Yet the mechanism of how…
cs.SI2025
When Life Paths Cross: Extracting Human Interactions in Time and Space from Wikipedia
Zhongyang Liu, Ying Zhang, Xiangyi Xiao +3
Interactions among notable individuals -- whether examined individually, in groups, or as networks -- often convey significant messages across cultural, economic, political, scient…