2 papers
cs.LG2026
Local-Order Auxiliary Losses Can Improve Autoencoder Reconstruction
Harvey Dam, Martin Burtscher, Tripti Agarwal +1
Mean-squared error is the default objective for training autoencoders, yet compressed reconstructions often depend not only on pointwise accuracy but also on preserving local spati…
cs.CL2025
Derailing Non-Answers via Logit Suppression at Output Subspace Boundaries in RLHF-Aligned Language Models
Harvey Dam, Jonas Knochelmann, Vinu Joseph +1
We introduce a method to reduce refusal rates of large language models (LLMs) on sensitive content without modifying model weights or prompts. Motivated by the observation that ref…