2 papers
cs.LG2026
IIB-LPO: Latent Policy Optimization via Iterative Information Bottleneck
Huilin Deng, Hongchen Luo, Yue Zhu +8
Recent advances in Reinforcement Learning with Verifiable Rewards (RLVR) for Large Language Model (LLM) reasoning have been hindered by a persistent challenge: exploration collapse…
cs.CY2025
Emergency Response Measures for Catastrophic AI Risk
James Zhang, Miles Kodama, Zongze Wu +3
Chinese authorities are extending the country's four-phase emergency response framework (prevent, warn, respond, and recover) to address risks from advanced artificial intelligence…