activity
20242026
collaborators

14 papers

cs.CV2026

HSA: Hierarchical Slot Attention for Multi-granularity Scene-Decomposition

Neelu Madan, Rongzhen Zhao, Andreas Mogelmose +4

Slot attention is a powerful framework for object-centric learning, decomposing visual scenes into latent slots through iterative competitive attention. However, existing methods s…

cs.CV2026

Internalizing Temporal Consistency in Video Object-Centric Learning without Explicit Regularization

Rongzhen Zhao, Zhiyuan Li, Juho Kannala +1

Video Object-Centric Learning (OCL) aims to represent objects as \textit{slot} vectors and maintain their consistency across frames. Slot-Slot Contrastive (SSC) loss has become the…

cs.CV2026

Cycle Consistency in Video Object-Centric Learning

Rongzhen Zhao, Zhiyuan Li, Ruonan Wei +2

Self-supervised video Object-Centric Learning (OCL) aims to discover distinct objects and associate them across time, whereas self-supervised Multi-Object Tracking (MOT) focuses on…

cs.CV2026

Object-Centric Vision Token Pruning for Vision Language Models

Guangyuan Li, Rongzhen Zhao, Jinhong Deng +2

In Vision Language Models (VLMs), vision tokens are quantity-heavy yet information-dispersed compared with language tokens, thus consume too much unnecessary computation. Pruning r…

cs.CV2026

Smoothing Slot Attention Iterations and Recurrences

Rongzhen Zhao, Wenyan Yang, Juho Kannala +1

Slot Attention (SA) lies at the heart of mainstream Object-Centric Learning (OCL). Image features can be aggregated into object-level representations by SA \textit{iteratively} ref…

cs.CV2026

Rethinking Temporal Consistency in Video Object-Centric Learning: From Prediction to Correspondence

Zhiyuan Li, Rongzhen Zhao, Wenyan Yang +3

The de facto approach in video object-centric learning maintains temporal consistency through learned dynamics modules that predict future object representations, called slots. We…