4 papers
The Distributed Detectability Band Against Marginal-Preserving Attacks
Zhang Qinqin, Gao Yuze
AI-control monitors score individual agent actions to detect misbehavior, but real harm can be distributed across many benign-looking steps, each individually below any per-step al…
A Pre-Registered Causal Partition of Self-Consistency Elicitation and Reward Design in RLVR
Yuze Gao
Reinforcement learning from verifiable rewards (RLVR) improves reasoning even when the reward signal is spurious -- assigning credit to the group-plurality answer rather than a gro…
Transferable and Undefendable Point Cloud Attacks via Medial Axis Transform
Keke Tang, Yuze Gao, Weilong Peng +3
Studying adversarial attacks on point clouds is essential for evaluating and improving the robustness of 3D deep learning models. However, most existing attack methods are develope…
SEACrowd: A Multilingual Multimodal Data Hub and Benchmark Suite for Southeast Asian Languages
Holy Lovenia, Rahmad Mahendra, Salsabil Maulana Akbar +58
Southeast Asia (SEA) is a region rich in linguistic diversity and cultural variety, with over 1,300 indigenous languages and a population of 671 million people. However, prevailing…