collaborators

31 papers

cs.LG2026

Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks

Nobin Sarwar, Shubhashis Roy Dipta, Zheyuan Liu +1

With the growing adoption of VLMs, DMs, LLMs, and AFMs, these multimodal foundation models can inadvertently encode sensitive, copyrighted, biased, or unsafe cross-modal associatio…

cs.LG2026

Multi-Way Representation Alignment

Akshit Achara, Tatiana Gaintseva, Mateo Mahaut +5

The Platonic Representation Hypothesis suggests that independently trained neural networks converge to increasingly similar latent spaces. However, current strategies for mapping t…

cs.AI2026

Subliminal Clocks: Latent Time Modelling in Diffusion Language Models

Maximo Rulli, Maximo Eduardo Rulli, Thomas Vaitses Fontanari +11

Diffusion Language Models (DLMs) have recently emerged as a promising alternative to autoregressive models. Unlike standard diffusion-based approaches, DLMs are not explicitly cond…

cs.LG2026

Demystifying Mergeability: Interpretable Properties to Predict Model Merging Success

Luca Zhou, Bo Zhao, Rose Yu +1

Model merging combines knowledge from separately fine-tuned models, yet the factors driving its success remain poorly understood. While recent work treats mergeability as an intrin…

cs.SD2026

PHALAR: Phasors for Learned Musical Audio Representations

Davide Marincione, Michele Mancusi, Giorgio Strano +4

Stem retrieval, the task of matching missing stems to a given audio submix, is a key challenge currently limited by models that discard temporal information. We introduce PHALAR, a…

cs.LG2026

TOAST: Transformer Optimization using Adaptive and Simple Transformations

Irene Cannistraci, Simone Antonelli, Emanuele Palumbo +4

Foundation models achieve state-of-the-art performance across different tasks, but their size and computational demands raise concerns about accessibility and sustainability. Exist…