94 citations · 228 across the 5 of their papers we have counts for
4 papers · 1 filter
FoodSense: A Multisensory Food Dataset and Benchmark for Predicting Taste, Smell, Texture, and Sound from Images
Sabab Ishraq, Aarushi Aarushi, Juncai Jiang +1
Humans routinely infer taste, smell, texture, and even sound from food images a phenomenon well studied in cognitive science. However, prior vision language research on food has fo…
Continental-Scale Building Detection from High Resolution Satellite Imagery
Wojciech Sirko, Sergii Kashubin, Marvin Ritter +7
Identifying the locations and footprints of buildings is vital for many practical and scientific purposes. Such information can be particularly useful in developing regions where a…
Representation learning from videos in-the-wild: An object-centric approach
Rob Romijnders, Aravindh Mahendran, Michael Tschannen +4
We propose a method to learn image representations from uncurated videos. We combine a supervised loss from off-the-shelf object detectors and self-supervised losses which naturall…
Self-Supervised Learning of Video-Induced Visual Invariances
Michael Tschannen, Josip Djolonga, Marvin Ritter +5
We propose a general framework for self-supervised learning of transferable visual representations based on Video-Induced Visual Invariances (VIVI). We consider the implicit hierar…