3 papers
cs.CL2026
Enhancing LLM Reasoning via Non-Human-Like Reasoning Path Preference Optimization
Junjie Lu, Yuliang Liu, Chaofeng Qu +4
Current approaches for strengthening LLM reasoning tend to introduce a training bias toward human-like reasoning trajectories. In step-wise preference optimization, in particular,…
q-fin.MF2025
Quantifying Bounded Rationality: Formal Verification of Simon's Satisficing Through Flexible Stochastic Dominance
Jingyuan Li, Zhou Lin
This paper introduces Flexible First-Order Stochastic Dominance (FFSD), a mathematically rigorous framework that formalizes Herbert Simon's concept of bounded rationality using the…
cs.AI2025
AdaptiveStep: Automatically Dividing Reasoning Step through Model Confidence
Yuliang Liu, Junjie Lu, Zhaoling Chen +10
Current approaches for training Process Reward Models (PRMs) often involve breaking down responses into multiple reasoning steps using rule-based techniques, such as using predefin…