3 papers
cs.CV2026
MobileSAM2: Lightweight Segment Anything for Spatial Intelligence
Kai Jiang, Jiaxing Huang, Jingyi Zhang +5
The recent large video foundation model, SAM2, enables segment anything in both images and videos, serving as a powerful base model for various applications. However, many of such…
cs.AI2025
AgentMath: Empowering Mathematical Reasoning for Large Language Models via Tool-Augmented Agent
Haipeng Luo, Huawen Feng, Qingfeng Sun +6
Large Reasoning Models (LRMs) like o3 and DeepSeek-R1 have achieved remarkable progress in reasoning tasks with long cot. However, they remain computationally inefficient and strug…
cs.CV2025
Towards Understanding Ambiguity Resolution in Multimodal Inference of Meaning
Yufei Wang, Adriana Kovashka, Loretta Fernández +2
We investigate a new setting for foreign language learning, where learners infer the meaning of unfamiliar words in a multimodal context of a sentence describing a paired image. We…