4 citations · 10 across the 6 of their papers we have counts for
6 papers
BRAVE: Broadening the visual encoding of vision-language models
Oğuzhan Fatih Kar, Alessio Tonioni, Petra Poklukar +3
Vision-language models (VLMs) are typically composed of a vision encoder, e.g. CLIP, and a language model (LM) that interprets the encoded features to solve downstream tasks. Despi…
Snap-it, Tap-it, Splat-it: Tactile-Informed 3D Gaussian Splatting for Reconstructing Challenging Surfaces
Mauro Comi, Alessio Tonioni, Max Yang +5
Touch and vision go hand in hand, mutually enhancing our ability to understand the world. From a research perspective, the problem of mixing touch and vision is underexplored and p…
InseRF: Text-Driven Generative Object Insertion in Neural 3D Scenes
Mohamad Shahbazi, Liesbeth Claessens, Michael Niemeyer +4
We introduce InseRF, a novel method for generative object insertion in the NeRF reconstructions of 3D scenes. Based on a user-provided textual description and a 2D bounding box in…
TextMesh: Generation of Realistic 3D Meshes From Text Prompts
Christina Tsalicoglou, Fabian Manhardt, Alessio Tonioni +2
The ability to generate highly realistic 2D images from mere text prompts has recently made huge progress in terms of speed and quality, thanks to the advent of image diffusion mod…
NeRF-Supervised Deep Stereo
Fabio Tosi, Alessio Tonioni, Daniele De Gregorio +1
We introduce a novel framework for training deep stereo networks effortlessly and without any ground-truth. By leveraging state-of-the-art neural rendering solutions, we generate s…
Learning Good Features to Transfer Across Tasks and Domains
Pierluigi Zama Ramirez, Adriano Cardace, Luca De Luigi +3
Availability of labelled data is the major obstacle to the deployment of deep learning algorithms for computer vision tasks in new domains. The fact that many frameworks adopted to…