10 citations · 10 across the 2 of their papers we have counts for
6 papers · 1 filter
SAMed-2: Selective Memory Enhanced Medical Segment Anything Model
Zhiling Yan, Sifan Song, Dingjie Song +11
Recent "segment anything" efforts show promise by learning from large-scale data, but adapting such models directly to medical images remains challenging due to the complexity of m…
EchoFM: Foundation Model for Generalizable Echocardiogram Analysis
Sekeun Kim, Pengfei Jin, Sifan Song +6
Foundation models have recently gained significant attention because of their generalizability and adaptability across multiple tasks and data distributions. Although medical found…
OracleSage: Towards Unified Visual-Linguistic Understanding of Oracle Bone Scripts through Cross-Modal Knowledge Fusion
Hanqi Jiang, Yi Pan, Junhao Chen +8
Oracle bone script (OBS), as China's earliest mature writing system, present significant challenges in automatic recognition due to their complex pictographic structures and diverg…
Examining the Commitments and Difficulties Inherent in Multimodal Foundation Models for Street View Imagery
Zhenyuan Yang, Xuhui Lin, Qinyi He +10
The emergence of Large Language Models (LLMs) and multimodal foundation models (FMs) has generated heightened interest in their applications that integrate vision and language. Thi…
Biomedical SAM 2: Segment Anything in Biomedical Images and Videos
Zhiling Yan, Weixiang Sun, Rong Zhou +8
Medical image segmentation and video object segmentation are essential for diagnosing and analyzing diseases by identifying and measuring biological structures. Recent advances in…
Eye-gaze Guided Multi-modal Alignment for Medical Representation Learning
Chong Ma, Hanqi Jiang, Wenting Chen +10
In the medical multi-modal frameworks, the alignment of cross-modality features presents a significant challenge. However, existing works have learned features that are implicitly…