2 papers
cs.CV2024
Similarity Guided Multimodal Fusion Transformer for Semantic Location Prediction in Social Media
Zhizhen Zhang, Ning Wang, Haojie Li +1
Semantic location prediction aims to derive meaningful location insights from multimodal social media posts, offering a more contextual understanding of daily activities than using…
cs.CV2024
Learning Pixel-wise Continuous Depth Representation via Clustering for Depth Completion
Chen Shenglun, Zhang Hong, Ma XinZhu +2
Depth completion is a long-standing challenge in computer vision, where classification-based methods have made tremendous progress in recent years. However, most existing classific…