3 papers
cs.LG2025
Probe-based Fine-tuning for Reducing Toxicity
Jan Wehner, Mario Fritz
Probes trained on model activations can detect undesirable behaviors like deception or biases that are difficult to identify from outputs alone. This makes them useful detectors to…
cs.LG2025
Taxonomy, Opportunities, and Challenges of Representation Engineering for Large Language Models
Jan Wehner, Sahar Abdelnabi, Daniel Tan +2
Representation Engineering (RepE) is a novel paradigm for controlling the behavior of LLMs. Unlike traditional approaches that modify inputs or fine-tune the model, RepE directly m…
cs.AI2025
Safety Must Precede the Deployment of Open-Ended AI
Ivaxi Sheth, Jan Wehner, Sahar Abdelnabi +2
AI advancements have been significantly driven by a combination of foundation models and curiosity-driven learning aimed at increasing capability and adaptability. Within this land…