3 papers
cs.CV2026
Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention
Daniel Shalam, Emanuel Ben Baruch, Avi Ben Cohen +1
Multimodal large language models can emit localized predictions, bounding boxes for objects and temporal windows for video and audio events, but they hallucinate these regions prol…
cs.CV2024
Unsupervised Representation Learning by Balanced Self Attention Matching
Daniel Shalam, Simon Korman
Many leading self-supervised methods for unsupervised representation learning, in particular those for embedding image features, are built on variants of the instance discriminatio…
cs.LG2024
The Balanced-Pairwise-Affinities Feature Transform
Daniel Shalam, Simon Korman
The Balanced-Pairwise-Affinities (BPA) feature transform is designed to upgrade the features of a set of input items to facilitate downstream matching or grouping related tasks. Th…