2 papers
cs.AI2026
Scaling the Scaling Logic: Agentic Meta-Synthesis of Logic Reasoning
Bowen Liu, Zhi Wu, Runquan Xie +2
Reinforcement Learning from Verifiable Rewards (RLVR) is bottlenecked by data: existing synthesis pipelines rely on expert-written code or fixed templates, confining growth to inst…
cs.CL2025
LLMs are Single-threaded Reasoners: Demystifying the Working Mechanism of Soft Thinking
Junhong Wu, Jinliang Lu, Zixuan Ren +4
Human cognition naturally engages with abstract and fluid concepts, whereas existing reasoning models often rely on generating discrete tokens, potentially constraining their expre…