4 papers
Self-Recognition Finetuning can Prevent and Reverse Emergent Misalignment
Arush Tagade, Shaoheng Zhou, Jiaxin Wen +1
Emergent misalignment (EM) has been linked to the activation of misaligned persona vectors and evil character traits, suggesting that EM operates through disruption of the model's…
LLM Assertiveness can be Mechanistically Decomposed into Emotional and Logical Components
Hikaru Tsujimura, Arush Tagade
Large Language Models (LLMs) often display overconfidence, presenting information with unwarranted certainty in high-stakes contexts. We investigate the internal basis of this beha…
Explaining Surface Layer Theory Departures in Marine Flux Profiles with Data-Driven Discovery
Jack Foxabbott, Leo Mckee-Reid, Andrew Cusick +8
Monin--Obukhov Similarity Theory (MOST), which underpins nearly all bulk estimates of surface fluxes in the atmospheric surface layer, assumes monotonic wind profiles and verticall…
Benchmarking the Discovery Engine
Jack Foxabbott, Arush Tagade, Andrew Cusick +6
The Discovery Engine is a general purpose automated system for scientific discovery, which combines machine learning with state-of-the-art ML interpretability to enable rapid and r…