3 papers
cs.LG2026
UBP2: Uncertainty-Balanced Preference Planning for Efficient Preference-based Reinforcement Learning
Mohamed Nabail, Leo Kaixuan Cheng, Jingmin Wang +1
Preference-based RL provides an approach to learning reward models from pairwise comparisons of behaviors, bypassing the need for explicit reward design. However, existing methods…
cs.SE2026
VulnAgent-R2: Evidence-Calibrated Multi-Agent Auditing for Repository-Level Vulnerability Detection
Renwei Meng, Haoyi Wu, Jingming Wang
Software vulnerabilities often depend on cross-file data flow, build options, framework conventions, and runtime guards, so isolated function classifiers produce fragile and poorly…
cs.AI2025
Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens
Jihwan Jeong, Xiaoyu Wang, Jingmin Wang +2
Offline reinforcement learning (RL) is crucial when online exploration is costly or unsafe but often struggles with high epistemic uncertainty due to limited data. Existing methods…