Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Neural Chameleons: Language Models Can Learn to Hide Their Thoughts from Unseen Activation Monitors
Max McGuinness, Alex Serrano, Luke Bailey +1
Activation monitoring, which probes a model's internal states using lightweight classifiers, is an emerging tool for AI safety. However, its worst-case robustness under a misalignm…
cs.LG2025
Path Integral Optimiser: Global Optimisation via Neural Schrödinger-Föllmer Diffusion
Max McGuinness, Eirik Fladmark, Francisco Vargas
We present an early investigation into the use of neural diffusion processes for global optimisation, focusing on Zhang et al.'s Path Integral Sampler. One can use the Boltzmann di…