Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention
Daniel Shalam, Emanuel Ben Baruch, Avi Ben Cohen +1
Multimodal large language models can emit localized predictions, bounding boxes for objects and temporal windows for video and audio events, but they hallucinate these regions prol…
cs.CV2024
Unsupervised Representation Learning by Balanced Self Attention Matching
Daniel Shalam, Simon Korman
Many leading self-supervised methods for unsupervised representation learning, in particular those for embedding image features, are built on variants of the instance discriminatio…