3 papers
cs.CL2026
A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models
Congmin Zheng, Jiachen Zhu, Zhuoying Ou +8
Although Large Language Models (LLMs) exhibit advanced reasoning ability, conventional alignment remains largely dominated by outcome reward models (ORMs) that judge only final ans…
stat.ML2025
Beyond Prior Limits: Addressing Distribution Misalignment in Particle Filtering
Yiwei Shi, Jingyu Hu, Yu Zhang +4
Particle filtering is a Bayesian inference method and a fundamental tool in state estimation for dynamic systems, but its effectiveness is often limited by the constraints of the i…
cs.LG2025
Belief-Contraction-Driven Active Inverse Source Localization and Characterization
Yiwei Shi, Mengyue Yang, Qi Zhang +3
Active inverse source localization and characterization (ISLC) in dynamic fields requires sequential decision making under partial observability, where a mobile sensor must infer l…