3 papers
cs.CV2026
Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention
Daniel Shalam, Emanuel Ben Baruch, Avi Ben Cohen +1
Multimodal large language models can emit localized predictions, bounding boxes for objects and temporal windows for video and audio events, but they hallucinate these regions prol…
cs.CV2025
LV-MAE: Learning Long Video Representations through Masked-Embedding Autoencoders
Ilan Naiman, Emanuel Ben-Baruch, Oron Anschel +4
In this work, we introduce long-video masked-embedding autoencoders (LV-MAE), a self-supervised learning framework for long video representation. Our approach treats short- and lon…
cs.CV2024
Distilling the Knowledge in Data Pruning
Emanuel Ben-Baruch, Adam Botach, Igor Kviatkovsky +2
With the increasing size of datasets used for training neural networks, data pruning becomes an attractive field of research. However, most current data pruning algorithms are limi…