3 papers
cs.LG2026
A Causal Model for Locating and Unlocking Sandbagging in Model Organisms
Hong Kiat Tan, Linh Le, David Williams-King
Sandbagging models strategically underperform on evaluations while retaining the capabilities being measured. The evaluations that guide frontier-model deployment and governance th…
cs.LG2026
Reference-Grafting Matches Fine-Tuning at Eliciting Sandbagged Capabilities
Linh Le, Hong Kiat Tan, David Williams-King
Sandbagging, in which a model deliberately underperforms on an evaluation despite retaining the underlying capability, threatens the safety evaluations that frontier-model governan…
cond-mat.mtrl-sci2026
PUDA: An AI-Native Hardware Harness for Self-Driving Laboratories
Zekun Ren, Hongzhao Tan, Jiaen Yee +1
Physical Unified Device Architecture (PUDA) is an AI-native hardware harness for self-driving laboratories (SDLs). Rather than building a human-centered graphical user interface (G…