12 papers
Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks
Nobin Sarwar, Shubhashis Roy Dipta, Zheyuan Liu +1
With the growing adoption of VLMs, DMs, LLMs, and AFMs, these multimodal foundation models can inadvertently encode sensitive, copyrighted, biased, or unsafe cross-modal associatio…
Robots Ask the Way: Communication-Enabled Social Navigation
Valentino Sacco, Luca Scofano, Indro Spinelli +1
Assistive autonomous robots operating in multi-agent environments require efficient strategies to locate specific individuals among multiple residents. Current social navigation me…
Bimanual Robot Manipulation via Multi-Agent In-Context Learning
Alessio Palma, Indro Spinelli, Vignesh Prasad +4
Language Models (LLMs) have emerged as powerful reasoning engines for embodied control. In particular, In-Context Learning (ICL) enables off-the-shelf, text-only LLMs to predict ro…
Quantifying Self-Preservation Bias in Large Language Models
Matteo Migliarini, Joaquin Pereira Pizzini, Luca Moresca +3
Instrumental convergence predicts that sufficiently advanced AI agents will resist shutdown, yet current safety training (RLHF) may obscure this risk by teaching models to deny sel…
Video Unlearning via Low-Rank Refusal Vector
Simone Facchiano, Stefano Saravalle, Matteo Migliarini +7
Video generative models achieve high-quality synthesis from natural-language prompts by leveraging large-scale web data. However, this training paradigm inherently exposes them to…
PhysTalk: Language-driven Real-time Physics in 3D Gaussian Scenes
Luca Collorone, Mert Kiray, Indro Spinelli +2
Realistic visual simulations are omnipresent, yet their creation requires computing time, rendering, and expert animation knowledge. Open-vocabulary visual effects generation from…