Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
IR: Contrastive Inverse Reinforcement Learning for Interpretable Detection and Mitigation of Reward Hacking
Mohammad Beigi, Ming Jin, Junshan Zhang +3
Reinforcement Learning from Human Feedback (RLHF) enables powerful LLM alignment but can introduce reward hacking - models exploit spurious correlations in proxy rewards without ge…
cs.AI2026
ChatAD: Reasoning-Enhanced Time-Series Anomaly Detection with Multi-Turn Instruction Evolution
Hui Sun, Chang Xu, Haonan Xie +7
LLM-driven Anomaly Detection (AD) helps enhance the understanding and explanatory abilities of anomalous behaviors in Time Series (TS). Existing methods face challenges of inadequa…