3 papers
cs.CV2025
S3Former: Self-supervised High-resolution Transformer for Solar PV Profiling
Minh Tran, Adrian De Luis, Haitao Liao +5
As the impact of climate change escalates, the global necessity to transition to sustainable energy sources becomes increasingly evident. Renewable energies have emerged as a viabl…
cs.CV2024
HENASY: Learning to Assemble Scene-Entities for Egocentric Video-Language Model
Khoa Vo, Thinh Phan, Kashu Yamazaki +2
Current video-language models (VLMs) rely extensively on instance-level alignment between video and language modalities, which presents two major limitations: (1) visual reasoning…
cs.CV2024
ShapeFormer: Shape Prior Visible-to-Amodal Transformer-based Amodal Instance Segmentation
Minh Tran, Winston Bounsavy, Khoa Vo +3
Amodal Instance Segmentation (AIS) presents a challenging task as it involves predicting both visible and occluded parts of objects within images. Existing AIS methods rely on a bi…