papers

Publications (17)

cs.CV2018

Discovery and usage of joint attention in images

Daniel Harari, Joshua B. Tenenbaum, Shimon Ullman

Joint visual attention is characterized by two or more individuals looking at a common target at the same time. The ability to identify joint attention in scenes, the people involv…

cs.CV2016

Do You See What I Mean? Visual Resolution of Linguistic Ambiguities

Yevgeni Berzak, Andrei Barbu, Daniel Harari +2

Understanding language goes hand in hand with the ability to integrate complex contextual information obtained via perception. In this work, we present a novel task for grounded la…

cs.CV2024

Why Not Use Your Textbook? Knowledge-Enhanced Procedure Planning of Instructional Videos

Kumaranage Ravindu Yasas Nagasinghe, Honglu Zhou, Malitha Gunawardhana +3

In this paper, we explore the capability of an agent to construct a logical sequence of action steps, thereby assembling a strategic procedural plan. This plan is crucial for navig…

cs.CV2026

Action Without Interaction: Probing the Physical Foundations of Video LMMs via Contact-Release Detection

Daniel Harari, Michael Sidorov, Chen Shterental +3

Large multi-modal models (LMMs) show increasing performance in realistic visual tasks for images and, more recently, for videos. For example, given a video sequence, such models ar…

cs.CV2026

Objects Before Words: Object-First Inductive Biases for Grounding Language in Child-View Video

Sathira Silva, Abrham Kahsay Gebreselasie, Muhammad Umer Sheikh +3

Learning grounded word meaning from natural experience requires resolving two ambiguities in infant-view recordings: when the named referent appears and where it is in a cluttered…

cs.CV2024

How Effective are Self-Supervised Models for Contact Identification in Videos

Malitha Gunawardhana, Limalka Sadith, Liel David +2

The exploration of video content via Self-Supervised Learning (SSL) models has unveiled a dynamic field of study, emphasizing both the complex challenges and unique opportunities i…