14 papers
HSA: Hierarchical Slot Attention for Multi-granularity Scene-Decomposition
Neelu Madan, Rongzhen Zhao, Andreas Mogelmose +4
Slot attention is a powerful framework for object-centric learning, decomposing visual scenes into latent slots through iterative competitive attention. However, existing methods s…
Internalizing Temporal Consistency in Video Object-Centric Learning without Explicit Regularization
Rongzhen Zhao, Zhiyuan Li, Juho Kannala +1
Video Object-Centric Learning (OCL) aims to represent objects as \textit{slot} vectors and maintain their consistency across frames. Slot-Slot Contrastive (SSC) loss has become the…
Cycle Consistency in Video Object-Centric Learning
Rongzhen Zhao, Zhiyuan Li, Ruonan Wei +2
Self-supervised video Object-Centric Learning (OCL) aims to discover distinct objects and associate them across time, whereas self-supervised Multi-Object Tracking (MOT) focuses on…
Object-Centric Vision Token Pruning for Vision Language Models
Guangyuan Li, Rongzhen Zhao, Jinhong Deng +2
In Vision Language Models (VLMs), vision tokens are quantity-heavy yet information-dispersed compared with language tokens, thus consume too much unnecessary computation. Pruning r…
Smoothing Slot Attention Iterations and Recurrences
Rongzhen Zhao, Wenyan Yang, Juho Kannala +1
Slot Attention (SA) lies at the heart of mainstream Object-Centric Learning (OCL). Image features can be aggregated into object-level representations by SA \textit{iteratively} ref…
Rethinking Temporal Consistency in Video Object-Centric Learning: From Prediction to Correspondence
Zhiyuan Li, Rongzhen Zhao, Wenyan Yang +3
The de facto approach in video object-centric learning maintains temporal consistency through learned dynamics modules that predict future object representations, called slots. We…