activity
20242026
collaborators

12 papers

cs.MM2026

AV-JEPA: Extending LeJEPA to Audio-Visual Self-Supervised Learning

Benjamin Robson, Santeri Mentu, Wenshuai Zhao +1

We present AV-JEPA, an elegant multimodal extension of LeJEPA to audio-visual self-supervised learning. Using an early-fusion Vision Transformer and modality dropout as masking, th…

cs.LG2026

Efficient Reinforcement Learning by Guiding World Models with Non-Curated Data

Yi Zhao, Aidan Scannell, Wenshuai Zhao +7

Leveraging offline data is a promising way to improve the sample efficiency of online reinforcement learning (RL). This paper expands the pool of usable data for offline-to-online…

cs.RO2026

Point Tracking Improves World Action Models

Jiarui Guan, Wenshuai Zhao, Yue Pei +3

Robot policy learning benefits from world-action models that capture environment dynamics, but pixel-level prediction entangles dynamics with nuisance factors such as lighting and…

cs.CV2026

Rethinking Temporal Consistency in Video Object-Centric Learning: From Prediction to Correspondence

Zhiyuan Li, Rongzhen Zhao, Wenyan Yang +3

The de facto approach in video object-centric learning maintains temporal consistency through learned dynamics modules that predict future object representations, called slots. We…

cs.RO2026

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing

Zhiyuan Li, Wenyan Yang, Wenshuai Zhao +4

Learning robotic manipulation from human videos is a promising solution to the data bottleneck in robotics, but the distribution shift between humans and robots remains a critical…

cs.AI2026

Closed-Loop Vision-Language Planning for Multi-Agent Coordination

Zhiyuan Li, Wenshuai Zhao, Joni Pajarinen

Cooperative multi-agent reinforcement learning (MARL) struggles with sample efficiency, interpretability, and generalization. While Large Language Models (LLMs) offer powerful plan…