2 papers
cs.CV2026
Action Without Interaction: Probing the Physical Foundations of Video LMMs via Contact-Release Detection
Daniel Harari, Michael Sidorov, Chen Shterental +3
Large multi-modal models (LMMs) show increasing performance in realistic visual tasks for images and, more recently, for videos. For example, given a video sequence, such models ar…
cs.CV2025
CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders
Shihab Aaqil Ahamed, Malitha Gunawardhana, Liel David +3
Current video-based Masked Autoencoders (MAEs) primarily focus on learning effective spatiotemporal representations from a visual perspective, which may lead the model to prioritiz…