3 papers
cs.CL2026
Explaining Intrinsic Moral Self-Correction with Mechanistic Interpretability
Yu-Ting Lee, Fu-Chieh Chang, Yu-En Shu +2
Intrinsic moral self-correction refers to the phenomenon where a language model refines its ethical judgments or aligns its outputs purely through prompting. While effective across…
cs.CL2025
Step-DeepResearch Technical Report
Chen Hu, Haikuo Du, Heng Wang +64
As LLMs shift toward autonomous agents, Deep Research has emerged as a pivotal metric. However, existing academic benchmarks like BrowseComp often fail to meet real-world demands f…
cs.AI2025
DialogueReason: Rule-Based RL Sparks Dialogue Reasoning in LLMs
Yubo Shu, Zhewei Huang, Xin Wu +3
We propose DialogueReason, a reasoning paradigm that uncovers the lost roles in monologue-style reasoning models, aiming to boost diversity and coherency of the reasoning process.…