5 citations · 26 across the 53 of their papers we have counts for
30 papers · 1 filter
GROM: Gradient-Free Rapid One-Shot Machine Unlearning
Paweł Batorski, Przemysław Spurek, Paul Swoboda
Machine unlearning has become a critical capability for safely removing specific, sensitive knowledge from large language models (LLMs). Current state-of-the-art approaches primari…
Spectral Energy Centroid: a Metric for Improving Performance and Analyzing Spectral Bias in Implicit Neural Representations
Tomasz Dądela, Adam Kania, Maciej Rut +1
Implicit Neural Representations (INRs) model continuous signals using multilayer perceptrons (MLPs), enabling compact, differentiable, and high-fidelity representations of data acr…
Stop Marginalizing My Dreams: Model Inversion via Laplace Kernel for Continual Learning
Patryk Krukowski, Jacek Tabor, Przemysław Spurek +2
Data-free continual learning (DFCIL) relies on model inversion to synthesize pseudo-samples and mitigate catastrophic forgetting. However, existing inversion methods are fundamenta…
SoftSAE: Dynamic Top-K Selection for Adaptive Sparse Autoencoders
Jakub Stępień, Marcin Mazur, Jacek Tabor +1
Sparse Autoencoders (SAEs) have become an important tool in mechanistic interpretability, helping to analyze internal representations in both Large Language Models (LLMs) and Visio…
REBEL: Hidden Knowledge Recovery via Evolutionary-Based Evaluation Loop
Patryk Rybak, Paweł Batorski, Paul Swoboda +1
Machine unlearning for LLMs aims to remove sensitive or copyrighted data from trained models. However, the true efficacy of current unlearning methods remains uncertain. Standard e…
Memory Self-Regeneration: Uncovering Hidden Knowledge in Unlearned Models
Agnieszka Polowczyk, Alicja Polowczyk, Joanna Waczyńska +2
The impressive capability of modern text-to-image models to generate realistic visuals has come with a serious drawback: they can be misused to create harmful, deceptive or unlawfu…