collaborators

13 papers

cs.MM2026

AV-JEPA: Extending LeJEPA to Audio-Visual Self-Supervised Learning

Benjamin Robson, Santeri Mentu, Wenshuai Zhao +1

We present AV-JEPA, an elegant multimodal extension of LeJEPA to audio-visual self-supervised learning. Using an early-fusion Vision Transformer and modality dropout as masking, th…

cs.RO2026

Point Tracking Improves World Action Models

Jiarui Guan, Wenshuai Zhao, Yue Pei +3

Robot policy learning benefits from world-action models that capture environment dynamics, but pixel-level prediction entangles dynamics with nuisance factors such as lighting and…

cs.CV2026

Cross-View Splatter: Feed-Forward View Synthesis with Georeferenced Images

Matias Turkulainen, Akshay Krishnan, Filippo Aleotti +6

We present Cross-View Splatter, a feed-forward method that predicts pixel-aligned Gaussian splats for outdoor scenes captured at ground level AND by satellite. Faithful reconstruct…

cs.LG2026

Privacy Leakage via Output Label Space and Differentially Private Continual Learning

Marlon Tobaben, Talal Alrawajfeh, Marcus Klasson +3

Differential privacy (DP) is a formal privacy framework that enables training machine learning (ML) models while protecting individuals' data. As pointed out by prior work, ML mode…

cs.CV2026

Latent-Compressed Variational Autoencoder for Video Diffusion Models

Jiarui Guan, Wenshuai Zhao, Zhengtao Zou +2

Video variational autoencoders (VAEs) used in latent diffusion models typically require a sufficiently large number of latent channels to ensure high-quality video reconstruction.…

cs.CV2026

PAWS: Perception of Articulation in the Wild at Scale from Egocentric Videos

Yihao Wang, Yang Miao, Wenshuai Zhao +8

Articulation perception aims to recover the motion and structure of articulated objects (e.g., drawers and cupboards), and is fundamental to 3D scene understanding in robotics, sim…