4 papers
Toward a Vision-Language Foundation Model for Medical Data: Multimodal Dataset and Benchmarks for Vietnamese PET/CT Report Generation
Huu Tien Nguyen, Dac Thai Nguyen, The Minh Duc Nguyen +11
Vision-Language Foundation Models (VLMs), trained on large-scale multimodal datasets, have driven significant advances in Artificial Intelligence (AI) by enabling rich cross-modal…
DualFit: A Two-Stage Virtual Try-On via Warping and Synthesis
Minh Tran, Johnmark Clements, Annie Prasanna +2
Virtual Try-On technology has garnered significant attention for its potential to transform the online fashion retail experience by allowing users to visualize how garments would l…
A2VIS: Amodal-Aware Approach to Video Instance Segmentation
Minh Tran, Thang Pham, Winston Bounsavy +2
Handling occlusion remains a significant challenge for video instance-level tasks like Multiple Object Tracking (MOT) and Video Instance Segmentation (VIS). In this paper, we propo…
Amodal Instance Segmentation with Diffusion Shape Prior Estimation
Minh Tran, Khoa Vo, Tri Nguyen +1
Amodal Instance Segmentation (AIS) presents an intriguing challenge, including the segmentation prediction of both visible and occluded parts of objects within images. Previous met…