6 papers
Results of the 1st Asynchronous CASTLE Challenge at the Joint Egocentric Vision Workshop in Conjunction with CVPR 2026
Luca Rossetto, Werner Bailer, Cathal Gurrin +5
This report summarizes the contributions and results of the 1st Asynchronous CASTLE Challenge at the Joint Egocentric Vision Workshop in conjunction with CVPR 2026.
The State-of-the-Art in Lifelog Retrieval: A Review of Progress at the ACM Lifelog Search Challenge Workshop 2022-24
Allie Tran, Werner Bailer, Duc-Tien Dang-Nguyen +8
The ACM Lifelog Search Challenge (LSC) is a venue that welcomes and compares systems that support the exploration of lifelog data, and in particular the retrieval of specific infor…
UNION: A Lightweight Target Representation for Efficient Zero-Shot Image-Guided Retrieval with Optional Textual Queries
Hoang-Bao Le, Allie Tran, Binh T. Nguyen +2
Image-Guided Retrieval with Optional Text (IGROT) is a general retrieval setting where a query consists of an anchor image, with or without accompanying text, aiming to retrieve se…
FIGROTD: A Friendly-to-Handle Dataset for Image Guided Retrieval with Optional Text
Hoang-Bao Le, Allie Tran, Binh T. Nguyen +2
Image-Guided Retrieval with Optional Text (IGROT) unifies visual retrieval (without text) and composed retrieval (with text). Despite its relevance in applications like Google Imag…
On the Brittleness of CLIP Text Encoders
Allie Tran, Luca Rossetto
Multimodal co-embedding models, especially CLIP, have advanced the state of the art in zero-shot classification and multimedia information retrieval in recent years by aligning ima…
The CASTLE 2024 Dataset: Advancing the Art of Multimodal Understanding
Luca Rossetto, Werner Bailer, Duc-Tien Dang-Nguyen +11
Egocentric video has seen increased interest in recent years, as it is used in a range of areas. However, most existing datasets are limited to a single perspective. In this paper,…