1 citations · 1 across the 3 of their papers we have counts for
3 papers
cs.CV2024
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization
Tan-Hanh Pham, Hoang-Nam Le, Phu-Vinh Nguyen +2
Visual Language Models have demonstrated remarkable capabilities across tasks, including visual question answering and image captioning. However, most models rely on text-based ins…
cs.CL2024
wav2graph: A Framework for Supervised Learning Knowledge Graph from Speech
Khai Le-Duc, Quy-Anh Dang, Tan-Hanh Pham +1
Knowledge graphs (KGs) enhance the performance of large language models (LLMs) and search engines by providing structured, interconnected data that improves reasoning and context-a…
eess.IV2024★ 1 cited
LiteGPT: Large Vision-Language Model for Joint Chest X-ray Localization and Classification Task
Khai Le-Duc, Ryan Zhang, Ngoc Son Nguyen +5
Vision-language models have been extensively explored across a wide range of tasks, achieving satisfactory performance; however, their application in medical imaging remains undere…