14 papers
Unsupervised Disentanglement Without Compromises : How Functional Orthogonality Enforces Identifiability
Mathieu Cyrille Simon, Pascal Frossard, Christophe De Vleeschouwer
This paper explores unsupervised disentangled representation learning from a functional perspective. We define latent concepts as factors that influence observations through locall…
Model soups need only one ingredient
Alireza Abdollahpoorrostam, Nikolaos Dimitriadis, Adam Hazimeh +1
Fine-tuning large pre-trained models on a target distribution often improves in-distribution (ID) accuracy, but at the cost of out-of-distribution (OOD) robustness as representatio…
Recursive Scaling in Masked Diffusion Models
Alba Carballo-Castro, Julianna Piskorz, Paulius Rauba +2
Masked diffusion models (MDMs) have recently emerged as a promising paradigm for sequence generation. Scaling MDMs is conventionally achieved by increasing the parameter count or t…
PriFT: Prior-Support Guided Supervised Fine-Tuning
Ke Wang, Shuangqi Li, Mathieu Salzmann +1
Supervised fine-tuning (SFT) is an efficient approach for downstream task adaptation and often serves as the initialization stage for reinforcement learning (RL), but it can show w…
Fixed-Point Masked Generative Modeling
Andrea Miele, Yiming Qin, Alba Carballo-Castro +2
Masked Generative Models (MGMs) enable parallel decoding and achieve strong performance across modalities, but require full-sequence bidirectional transformers at every step, makin…
MEMOIR: Lifelong Model Editing with Minimal Overwrite and Informed Retention for LLMs
Ke Wang, Yiming Qin, Nikolaos Dimitriadis +2
Language models deployed in real-world systems often require post-hoc updates to incorporate new or corrected knowledge. However, editing such models efficiently and reliably-witho…