activity
20202026
most citedOptimistic Multi-Agent Policy Gradient

1 citations · 4 across the 17 of their papers we have counts for

collaborators

18 papers

cs.CV2026

Temporal Self-Distillation: Learning Visual State Tracking in Videos Without Supervision

Shravan Venkatraman, Wenshuai Zhao, Mohammad Hassan Vali +1

We introduce ST (Self-Supervised Self-Distillation over Time), which, to the best of our knowledge, is the first fully self-contained framework for continuous video state track…

cs.MM2026

AV-JEPA: Extending LeJEPA to Audio-Visual Self-Supervised Learning

Benjamin Robson, Santeri Mentu, Wenshuai Zhao +1

We present AV-JEPA, an elegant multimodal extension of LeJEPA to audio-visual self-supervised learning. Using an early-fusion Vision Transformer and modality dropout as masking, th…

cs.RO2026

Point Tracking Improves World Action Models

Jiarui Guan, Wenshuai Zhao, Yue Pei +3

Robot policy learning benefits from world-action models that capture environment dynamics, but pixel-level prediction entangles dynamics with nuisance factors such as lighting and…

cs.CV2026

Rethinking Temporal Consistency in Video Object-Centric Learning: From Prediction to Correspondence

Zhiyuan Li, Rongzhen Zhao, Wenyan Yang +3

The de facto approach in video object-centric learning maintains temporal consistency through learned dynamics modules that predict future object representations, called slots. We…

cs.RO2026

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing

Zhiyuan Li, Wenyan Yang, Wenshuai Zhao +4

Learning robotic manipulation from human videos is a promising solution to the data bottleneck in robotics, but the distribution shift between humans and robots remains a critical…

cs.CV2026

Latent-Compressed Variational Autoencoder for Video Diffusion Models

Jiarui Guan, Wenshuai Zhao, Zhengtao Zou +2

Video variational autoencoders (VAEs) used in latent diffusion models typically require a sufficiently large number of latent channels to ensure high-quality video reconstruction.…