1 citations · 3 across the 11 of their papers we have counts for
11 papers
Internalizing Temporal Consistency in Video Object-Centric Learning without Explicit Regularization
Rongzhen Zhao, Zhiyuan Li, Juho Kannala +1
Video Object-Centric Learning (OCL) aims to represent objects as \textit{slot} vectors and maintain their consistency across frames. Slot-Slot Contrastive (SSC) loss has become the…
Cycle Consistency in Video Object-Centric Learning
Rongzhen Zhao, Zhiyuan Li, Ruonan Wei +2
Self-supervised video Object-Centric Learning (OCL) aims to discover distinct objects and associate them across time, whereas self-supervised Multi-Object Tracking (MOT) focuses on…
Rethinking Temporal Consistency in Video Object-Centric Learning: From Prediction to Correspondence
Zhiyuan Li, Rongzhen Zhao, Wenyan Yang +3
The de facto approach in video object-centric learning maintains temporal consistency through learned dynamics modules that predict future object representations, called slots. We…
Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing
Zhiyuan Li, Wenyan Yang, Wenshuai Zhao +4
Learning robotic manipulation from human videos is a promising solution to the data bottleneck in robotics, but the distribution shift between humans and robots remains a critical…
Sparsely Supervised Diffusion
Wenshuai Zhao, Zhiyuan Li, Yi Zhao +5
Diffusion models have shown remarkable success across a wide range of generative tasks. However, they often suffer from spatially inconsistent generation, arguably due to the inher…
Closed-Loop Vision-Language Planning for Multi-Agent Coordination
Zhiyuan Li, Wenshuai Zhao, Joni Pajarinen
Cooperative multi-agent reinforcement learning (MARL) struggles with sample efficiency, interpretability, and generalization. While Large Language Models (LLMs) offer powerful plan…