2 papers
cs.CR2026
Proof-of-Guardrail in AI Agents and What (Not) to Trust from It
Xisen Jin, Michael Duan, Qin Lin +4
As AI agents become widely deployed as online services, users often rely on an agent developer's claim about how safety is enforced, which introduces a threat where safety measures…
cs.CR2026
LATTICE: Evaluating Decision Support Utility of Crypto Agents
Aaron Chan, Tengfei Li, Tianyi Xiao +3
We introduce LATTICE, a benchmark for evaluating the decision support utility of crypto agents in realistic user-facing scenarios. Prior crypto agent benchmarks mainly focus on rea…