From the 1 of 6 linked papers with an AI index.
2 citations · 2 across the 4 of their papers we have counts for
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Policy of Thoughts: Scaling Test-Time Training for LLM Reasoning via Online Policy Evolution
Zhengbo Jiao, Hongyu Xian, Qinglong Wang +5
The paper introduces Policy of Thoughts (PoT), a test‑time training framework that continuously updates a lightweight LoRA adapter using online policy optimization to improve large…
cs.AI2026
DeepLook: Deeper Thinking with Lookahead
Tingxin Yang, Zefeng Wang, Mengyue Wang +2
Inference-time scaling has emerged as a powerful paradigm for improving large language model reasoning, often delivering larger gains on difficult reasoning tasks than parameter sc…