Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
PhyVLLM: Physics-Guided Video Language Model with Motion-Appearance Disentanglement
Yu-Wei Zhan, Xin Wang, Hong Chen +6
Video Large Language Models (Video LLMs) have shown impressive performance across a wide range of video-language tasks. However, they often fail in scenarios requiring a deeper und…
cs.CV2024
Multi-weather Cross-view Geo-localization Using Denoising Diffusion Models
Tongtong Feng, Qing Li, Xin Wang +3
Cross-view geo-localization in GNSS-denied environments aims to determine an unknown location by matching drone-view images with the correct geo-tagged satellite-view images from a…