2 papers
cs.CV2026
R3D: Quantitative 3D Spatial Reasoning for Egocentric Wearables
Maxwell Horton, Wei Lu, Quan Tran +6
Quantitative 3D spatial reasoning from egocentric RGB-D video is a critical capability for next-generation wearable assistants. Yet existing benchmarks do not reflect the challenge…
cs.IR2022
Normalized Contrastive Learning for Text-Video Retrieval
Yookoon Park, Mahmoud Azab, Bo Xiong +4
Cross-modal contrastive learning has led the recent advances in multimodal retrieval with its simplicity and effectiveness. In this work, however, we reveal that cross-modal contra…