1 citations · 4 across the 17 of their papers we have counts for
18 papers
Temporal Self-Distillation: Learning Visual State Tracking in Videos Without Supervision
Shravan Venkatraman, Wenshuai Zhao, Mohammad Hassan Vali +1
We introduce ST (Self-Supervised Self-Distillation over Time), which, to the best of our knowledge, is the first fully self-contained framework for continuous video state track…
AV-JEPA: Extending LeJEPA to Audio-Visual Self-Supervised Learning
Benjamin Robson, Santeri Mentu, Wenshuai Zhao +1
We present AV-JEPA, an elegant multimodal extension of LeJEPA to audio-visual self-supervised learning. Using an early-fusion Vision Transformer and modality dropout as masking, th…
Point Tracking Improves World Action Models
Jiarui Guan, Wenshuai Zhao, Yue Pei +3
Robot policy learning benefits from world-action models that capture environment dynamics, but pixel-level prediction entangles dynamics with nuisance factors such as lighting and…
Rethinking Temporal Consistency in Video Object-Centric Learning: From Prediction to Correspondence
Zhiyuan Li, Rongzhen Zhao, Wenyan Yang +3
The de facto approach in video object-centric learning maintains temporal consistency through learned dynamics modules that predict future object representations, called slots. We…
Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing
Zhiyuan Li, Wenyan Yang, Wenshuai Zhao +4
Learning robotic manipulation from human videos is a promising solution to the data bottleneck in robotics, but the distribution shift between humans and robots remains a critical…
Latent-Compressed Variational Autoencoder for Video Diffusion Models
Jiarui Guan, Wenshuai Zhao, Zhengtao Zou +2
Video variational autoencoders (VAEs) used in latent diffusion models typically require a sufficiently large number of latent channels to ensure high-quality video reconstruction.…