4 papers
Safety Must Precede the Deployment of Open-Ended AI
Ivaxi Sheth, Jan Wehner, Sahar Abdelnabi +2
AI advancements have been significantly driven by a combination of foundation models and curiosity-driven learning aimed at increasing capability and adaptability. Within this land…
Probe-based Fine-tuning for Reducing Toxicity
Jan Wehner, Mario Fritz
Probes trained on model activations can detect undesirable behaviors like deception or biases that are difficult to identify from outputs alone. This makes them useful detectors to…
Taxonomy, Opportunities, and Challenges of Representation Engineering for Large Language Models
Jan Wehner, Sahar Abdelnabi, Daniel Tan +2
Representation Engineering (RepE) is a novel paradigm for controlling the behavior of LLMs. Unlike traditional approaches that modify inputs or fine-tune the model, RepE directly m…
Explaining Learned Reward Functions with Counterfactual Trajectories
Jan Wehner, Frans Oliehoek, Luciano Cavalcante Siebert
Learning rewards from human behaviour or feedback is a promising approach to aligning AI systems with human values but fails to consistently extract correct reward functions. Inter…