13 citations · 22 across the 18 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
TI-JEPA: An Innovative Energy-based Joint Embedding Strategy for Text-Image Multimodal Systems
Khang H. N. Vo, Duc P. T. Nguyen, Thong Nguyen +1
This paper focuses on multimodal alignment within the realm of Artificial Intelligence, particularly in text and image modalities. The semantic gap between the textual and visual m…
cs.CV2024
ViConsFormer: Constituting Meaningful Phrases of Scene Texts using Transformer-based Method in Vietnamese Text-based Visual Question Answering
Nghia Hieu Nguyen, Tho Thanh Quan, Ngan Luu-Thuy Nguyen
Text-based VQA is a challenging task that requires machines to use scene texts in given images to yield the most appropriate answer for the given question. The main challenge of te…