6 papers
HSA: Hierarchical Slot Attention for Multi-granularity Scene-Decomposition
Neelu Madan, Rongzhen Zhao, Andreas Mogelmose +4
Slot attention is a powerful framework for object-centric learning, decomposing visual scenes into latent slots through iterative competitive attention. However, existing methods s…
A Hyperbolic Perspective on Hierarchical Structure in Object-Centric Scene Representations
Neelu Madan, Ãlex Pujol, Andreas Møgelmose +4
Slot attention has emerged as a powerful framework for unsupervised object-centric learning, decomposing visual scenes into a small set of compact vector representations called \em…
Self-Distillation of Hidden Layers for Self-Supervised Representation Learning
Scott C. Lowe, Anthony Fuller, Sageev Oore +2
The landscape of self-supervised learning (SSL) is currently dominated by generative approaches (e.g. MAE) that reconstruct raw low-level data, and predictive approaches (e.g. I-JE…
A multi-modal dataset for insect biodiversity with imagery and DNA at the trap and individual level
Johanna Orsholm, John Quinto, Hannu Autto +26
Insects comprise millions of species, many experiencing severe population declines under environmental and habitat changes. High-throughput approaches are crucial for accelerating…
LoSA: Long-Short-range Adapter for Scaling End-to-End Temporal Action Localization
Akshita Gupta, Gaurav Mittal, Ahmed Magooda +3
Temporal Action Localization (TAL) involves localizing and classifying action snippets in an untrimmed video. The emergence of large video foundation models has led RGB-only video…
Open-Vocabulary Temporal Action Localization using Multimodal Guidance
Akshita Gupta, Aditya Arora, Sanath Narayan +3
Open-Vocabulary Temporal Action Localization (OVTAL) enables a model to recognize any desired action category in videos without the need to explicitly curate training data for all…