Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
S-SPPO: Semantic-Calibrated Self-Play Preference Optimization
Xiwen Chen, Wenhui Zhu, Jingjing Wang +13
Aligning Large Language Models (LLMs) with human preferences is often formulated via Direct Preference Optimization (DPO). However, the standard Bradley-Terry instantiation of DPO…
cs.AI2025
Anomalous Decision Discovery using Inverse Reinforcement Learning
Ashish Bastola, Mert D. Pesé, Long Cheng +2
Anomaly detection plays a critical role in Autonomous Vehicles (AVs) by identifying unusual behaviors through perception systems that could compromise safety and lead to hazardous…