Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
MobileSAM2: Lightweight Segment Anything for Spatial Intelligence
Kai Jiang, Jiaxing Huang, Jingyi Zhang +5
The recent large video foundation model, SAM2, enables segment anything in both images and videos, serving as a powerful base model for various applications. However, many of such…
cs.CV2025
Towards Understanding Ambiguity Resolution in Multimodal Inference of Meaning
Yufei Wang, Adriana Kovashka, Loretta Fernández +2
We investigate a new setting for foreign language learning, where learners infer the meaning of unfamiliar words in a multimodal context of a sentence describing a paired image. We…