5 papers · 1 filter
GRASP: Learning to Ground Social Reasoning in Multi-Person Non-Verbal Interactions
Junho Kim, Xu Cao, Houze Yang +6
Understanding social interactions requires reasoning over subtle non-verbal cues, yet current multimodal large language models (MLLMs) often fail to identify who interacts with who…
MEBench: A Novel Benchmark for Understanding Mutual Exclusivity Bias in Vision-Language Models
Anh Thai, Stefan Stojanov, Zixuan Huang +2
This paper introduces MEBench, a novel benchmark for evaluating mutual exclusivity (ME) bias, a cognitive phenomenon observed in children during word learning. Unlike traditional M…
Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives
Kristen Grauman, Andrew Westbury, Lorenzo Torresani +98
We present Ego-Exo4D, a diverse, large-scale multimodal multiview video dataset and benchmark challenge. Ego-Exo4D centers around simultaneously-captured egocentric and exocentric…
Leveraging Object Priors for Point Tracking
Bikram Boote, Anh Thai, Wenqi Jia +4
Point tracking is a fundamental problem in computer vision with numerous applications in AR and robotics. A common failure mode in long-term point tracking occurs when the predicte…
Modeling Multimodal Social Interactions: New Challenges and Baselines with Densely Aligned Representations
Sangmin Lee, Bolin Lai, Fiona Ryan +2
Understanding social interactions involving both verbal and non-verbal cues is essential for effectively interpreting social situations. However, most prior works on multimodal soc…