3 papers
cs.CR2026
The Distributed Detectability Band Against Marginal-Preserving Attacks
Zhang Qinqin, Gao Yuze
AI-control monitors score individual agent actions to detect misbehavior, but real harm can be distributed across many benign-looking steps, each individually below any per-step al…
cs.AI2026
A Pre-Registered Causal Partition of Self-Consistency Elicitation and Reward Design in RLVR
Yuze Gao
Reinforcement learning from verifiable rewards (RLVR) improves reasoning even when the reward signal is spurious -- assigning credit to the group-plurality answer rather than a gro…
cs.CV2025
Transferable and Undefendable Point Cloud Attacks via Medial Axis Transform
Keke Tang, Yuze Gao, Weilong Peng +3
Studying adversarial attacks on point clouds is essential for evaluating and improving the robustness of 3D deep learning models. However, most existing attack methods are develope…