5 papers
DualFit: A Two-Stage Virtual Try-On via Warping and Synthesis
Minh Tran, Johnmark Clements, Annie Prasanna +2
Virtual Try-On technology has garnered significant attention for its potential to transform the online fashion retail experience by allowing users to visualize how garments would l…
S3Former: Self-supervised High-resolution Transformer for Solar PV Profiling
Minh Tran, Adrian De Luis, Haitao Liao +5
As the impact of climate change escalates, the global necessity to transition to sustainable energy sources becomes increasingly evident. Renewable energies have emerged as a viabl…
A2VIS: Amodal-Aware Approach to Video Instance Segmentation
Minh Tran, Thang Pham, Winston Bounsavy +2
Handling occlusion remains a significant challenge for video instance-level tasks like Multiple Object Tracking (MOT) and Video Instance Segmentation (VIS). In this paper, we propo…
HENASY: Learning to Assemble Scene-Entities for Egocentric Video-Language Model
Khoa Vo, Thinh Phan, Kashu Yamazaki +2
Current video-language models (VLMs) rely extensively on instance-level alignment between video and language modalities, which presents two major limitations: (1) visual reasoning…
Amodal Instance Segmentation with Diffusion Shape Prior Estimation
Minh Tran, Khoa Vo, Tri Nguyen +1
Amodal Instance Segmentation (AIS) presents an intriguing challenge, including the segmentation prediction of both visible and occluded parts of objects within images. Previous met…