3 citations · 6 across the 10 of their papers we have counts for
3 papers · 1 filter
Cross-domain Few-shot Object Detection with Multi-modal Textual Enrichment
Zeyu Shangguan, Daniel Seita, Mohammad Rostami
Advancements in cross-modal feature extraction and integration have significantly enhanced performance in few-shot learning tasks. However, current multi-modal object detection (MM…
PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding
Wei Chow, Jiageng Mao, Boyi Li +3
Understanding the physical world is a fundamental challenge in embodied AI, critical for enabling agents to perform complex tasks and operate safely in real-world environments. Whi…
Cross-domain Multi-modal Few-shot Object Detection via Rich Text
Zeyu Shangguan, Daniel Seita, Mohammad Rostami
Cross-modal feature extraction and integration have led to steady performance improvements in few-shot learning tasks due to generating richer features. However, existing multi-mod…