2 papers
cs.CV2026
HANCLIP: A Family of Hyperbolic Angular Negation Vision Language Models
Hoang-Bao Le, Aiden Durrant, Thai Son Mai +3
Vision-Language Models (VLMs) are typically pre-trained on large-scale image-text datasets to capture semantic correspondences between visual content and natural language. However,…
cs.LG2024
ExGra-Med: Extended Context Graph Alignment for Medical Vision-Language Models
Duy M. H. Nguyen, Nghiem T. Diep, Trung Q. Nguyen +10
State-of-the-art medical multi-modal LLMs (med-MLLMs), such as LLaVA-Med and BioMedGPT, primarily depend on scaling model size and data volume, with training driven largely by auto…