Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Failing to See or Failing to Know? Attributing Errors in Vision-Language Models
Khang Nhat Hoang Vo, Artem Vazhentsev, Artem Shelmanov +2
Vision-language models (VLMs) can recognize entities in clear images yet still fail when answering questions that require factual knowledge beyond what is directly observable. Prio…
cs.CV2025
TI-JEPA: An Innovative Energy-based Joint Embedding Strategy for Text-Image Multimodal Systems
Khang H. N. Vo, Duc P. T. Nguyen, Thong Nguyen +1
This paper focuses on multimodal alignment within the realm of Artificial Intelligence, particularly in text and image modalities. The semantic gap between the textual and visual m…