From the 1 of 20 linked papers with an AI index.
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
ReasoningShield: Safety Detection over Reasoning Traces of Large Reasoning Models
Changyi Li, Jiayi Wang, Xudong Pan +2
Large Reasoning Models (LRMs) leverage transparent reasoning traces, known as Chain-of-Thoughts (CoTs), to break down complex problems into intermediate steps and derive final answ…
cs.CL2024
Frontier AI systems have surpassed the self-replicating red line
Xudong Pan, Jiarun Dai, Yihe Fan +1
Successful self-replication under no human assistance is the essential step for AI to outsmart the human beings, and is an early signal for rogue AIs. That is why self-replication…