Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Taxonomy, Opportunities, and Challenges of Representation Engineering for Large Language Models
Jan Wehner, Sahar Abdelnabi, Daniel Tan +2
Representation Engineering (RepE) is a novel paradigm for controlling the behavior of LLMs. Unlike traditional approaches that modify inputs or fine-tune the model, RepE directly m…
cs.LG2023
Hazards from Increasingly Accessible Fine-Tuning of Downloadable Foundation Models
Alan Chan, Ben Bucknall, Herbie Bradley +1
Public release of the weights of pretrained foundation models, otherwise known as downloadable access \citep{solaiman_gradient_2023}, enables fine-tuning without the prohibitive ex…