2 papers
cs.AI2026
Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks
Youting Wang, Xiao Han, Dingyan Shang +2
Agent-safety benchmarks measure different behaviors, and their scores get quoted interchangeably as an agent's safety. We treat four of them (R-Judge, InjecAgent, AgentHarm, AgentD…
cs.LG2026
When LLM Reward Design Fails: Diagnostic-Driven Refinement for Sparse Structured RL
Youting Wang, Yuan Tang, Bowen Liu +2
For sparse, structured reinforcement-learning tasks with semantic reward-function interfaces, LLM-generated reward shaping is better framed as debugging than one-shot generation. W…