adaptive inference 1mathematical reasoning 1reinforcement learning 1self-verification 1test-time compute 1
From the 1 of 24 linked papers with an AI index.
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute
Hongyu Chen, Liang Lin, Guangrun Wang
The paper proposes Self‑Verifying Refinement (SVR), a reinforcement‑learning framework that lets language models decide when to stop refining answers by using their own correctness…
cs.AI2026
OOWM: Structuring Embodied Reasoning and Planning via Object-Oriented Programmatic World Modeling
Hongyu Chen, Liang Lin, Guangrun Wang
Standard Chain-of-Thought (CoT) prompting empowers Large Language Models (LLMs) with reasoning capabilities, yet its reliance on linear natural language is inherently insufficient…