Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute
Hongyu Chen, Liang Lin, Guangrun Wang
Scaling test-time computation can improve language-model reasoning, but uniform budgets waste computation on easy inputs, while verifier-guided refinement relies on external feedba…
cs.AI2026
OOWM: Structuring Embodied Reasoning and Planning via Object-Oriented Programmatic World Modeling
Hongyu Chen, Liang Lin, Guangrun Wang
Standard Chain-of-Thought (CoT) prompting empowers Large Language Models (LLMs) with reasoning capabilities, yet its reliance on linear natural language is inherently insufficient…