activity
20242026
most citedOrganizing Unstructured Image Collections using Natural Language

1 citations · 1 across the 3 of their papers we have counts for

collaborators

6 papers

cs.CV2026

Lightweight 3D Feature Pretraining by Bayesian Inversion of 2D Foundation Models

Marwane Hariat, Gianni Franchi, David Filliat +1

We present Casper3D, a lightweight probabilistic framework for converting noisy multi-view 2D foundation-model embeddings into a latent 3D semantic representation. We model view-le…

cs.CV2026

Leveraging Visual Signals for Robust Token-Level Uncertainty in Vision-Language Generation

Joseph Hoche, David Brellmann, Gianni Franchi

Uncertainty quantification (UQ) remains a critical challenge in Large Vision Language Models (LVLMs) for reliable predictions and real-world deployment. However, most existing meth…

cs.CV20261 cited

Organizing Unstructured Image Collections using Natural Language

Mingxuan Liu, Zhun Zhong, Jun Li +3

In this work, we introduce and study the novel task of Open-ended Semantic Multiple Clustering (OpenSMC). Given a large, unstructured image collection, the goal is to automatically…

cs.AI2025

Collaborative Instance Object Navigation: Leveraging Uncertainty-Awareness to Minimize Human-Agent Dialogues

Francesco Taioli, Edoardo Zorzi, Gianni Franchi +4

Language-driven instance object navigation assumes that human users initiate the task by providing a detailed description of the target instance to the embodied agent. While this d…

cs.CV2025

Hierarchical Light Transformer Ensembles for Multimodal Trajectory Forecasting

Adrien Lafage, Mathieu Barbier, Gianni Franchi +1

Accurate trajectory forecasting is crucial for the performance of various systems, such as advanced driver-assistance systems and self-driving vehicles. These forecasts allow us to…

cs.CV2024

Frustratingly Easy Test-Time Adaptation of Vision-Language Models

Matteo Farina, Gianni Franchi, Giovanni Iacca +2

Vision-Language Models seamlessly discriminate among arbitrary semantic categories, yet they still suffer from poor generalization when presented with challenging examples. For thi…