3 papers
cs.LG2026
Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility
Mohsen Hariri, Weicong Chen, Nahal Shahini +11
Large language models can solve harder reasoning problems with more inference-time compute. The term "test-time scaling," however, covers several inference algorithms: extending de…
cs.RO2026
Imagining Recovery: Inference-Time Counterfactual Realignment for Vision-Language-Action Models
Yanyan Zhang, Disheng Liu, Kai Ye +6
Vision-language-action (VLA) models have improved the flexibility and generality of robotic manipulation, yet they remain fragile to online disruptions, such as changes in task goa…
cs.AI2026
STRIDE: Strategic Trajectory Reasoning via Discriminative Estimation for Verifiable Reinforcement Learning
Qinjian Zhao, Zhihao Dou, Dinggen Zhang +10
Reinforcement Learning with Verifiable Rewards (RLVR) has become an effective post-training paradigm for improving the reasoning abilities of large language models. However, existi…