13 papers
AV-JEPA: Extending LeJEPA to Audio-Visual Self-Supervised Learning
Benjamin Robson, Santeri Mentu, Wenshuai Zhao +1
We present AV-JEPA, an elegant multimodal extension of LeJEPA to audio-visual self-supervised learning. Using an early-fusion Vision Transformer and modality dropout as masking, th…
Point Tracking Improves World Action Models
Jiarui Guan, Wenshuai Zhao, Yue Pei +3
Robot policy learning benefits from world-action models that capture environment dynamics, but pixel-level prediction entangles dynamics with nuisance factors such as lighting and…
Cross-View Splatter: Feed-Forward View Synthesis with Georeferenced Images
Matias Turkulainen, Akshay Krishnan, Filippo Aleotti +6
We present Cross-View Splatter, a feed-forward method that predicts pixel-aligned Gaussian splats for outdoor scenes captured at ground level AND by satellite. Faithful reconstruct…
Privacy Leakage via Output Label Space and Differentially Private Continual Learning
Marlon Tobaben, Talal Alrawajfeh, Marcus Klasson +3
Differential privacy (DP) is a formal privacy framework that enables training machine learning (ML) models while protecting individuals' data. As pointed out by prior work, ML mode…
Latent-Compressed Variational Autoencoder for Video Diffusion Models
Jiarui Guan, Wenshuai Zhao, Zhengtao Zou +2
Video variational autoencoders (VAEs) used in latent diffusion models typically require a sufficiently large number of latent channels to ensure high-quality video reconstruction.…
PAWS: Perception of Articulation in the Wild at Scale from Egocentric Videos
Yihao Wang, Yang Miao, Wenshuai Zhao +8
Articulation perception aims to recover the motion and structure of articulated objects (e.g., drawers and cupboards), and is fundamental to 3D scene understanding in robotics, sim…