1 citations · 4 across the 12 of their papers we have counts for
6 papers · 1 filter
NeRF-Insert: 3D Local Editing with Multimodal Control Signals
Benet Oriol Sabat, Alessandro Achille, Matthew Trager +1
We propose NeRF-Insert, a NeRF editing framework that allows users to make high-quality local edits with a flexible level of control. Unlike previous work that relied on image-to-i…
Multi-Modal Hallucination Control by Visual Information Grounding
Alessandro Favero, Luca Zancato, Matthew Trager +5
Generative Vision-Language Models (VLMs) are prone to generate plausible-sounding textual answers that, however, are not always grounded in the input image. We investigate this phe…
Interpretable Measures of Conceptual Similarity by Complexity-Constrained Descriptive Auto-Encoding
Alessandro Achille, Greg Ver Steeg, Tian Yu Liu +3
Quantifying the degree of similarity between images is a key copyright issue for image-based machine learning. In legal doctrine however, determining the degree of similarity betwe…
Towards Visual Foundational Models of Physical Scenes
Chethan Parameshwara, Alessandro Achille, Matthew Trager +7
We describe a first step towards learning general-purpose visual representations of physical scenes using only image prediction as a training criterion. To do so, we first define "…
Prompt Algebra for Task Composition
Pramuditha Perera, Matthew Trager, Luca Zancato +2
We investigate whether prompts learned independently for different tasks can be later combined through prompt algebra to obtain a model that supports composition of tasks. We consi…
Train/Test-Time Adaptation with Retrieval
Luca Zancato, Alessandro Achille, Tian Yu Liu +3
We introduce Train/Test-Time Adaptation with Retrieval (), a method to adapt models both at train and test time by means of a retrieval module and a searchable pool of…