activity
20242026
most citedMulti-weather Cross-view Geo-localization Using Denoising Diffusion Models

22 citations · 22 across the 3 of their papers we have counts for

collaborators

5 papers

cs.CV2026

AV-Unified: A Unified Framework for Audio-visual Scene Understanding

Guangyao Li, Xin Wang, Wenwu Zhu

When humans perceive the world, they naturally integrate multiple audio-visual tasks within dynamic, real-world scenes. However, current works such as event localization, parsing,…

cs.CV2026

Think with Grounding: Curriculum Reinforced Reasoning with Video Grounding for Long Video Understanding

Houlun Chen, Xin Wang, Guangyao Li +4

Long video understanding is challenging due to rich and complicated multimodal clues in long temporal range.Current methods adopt reasoning to improve the model's ability to analyz…

cs.CV2025

PhyVLLM: Physics-Guided Video Language Model with Motion-Appearance Disentanglement

Yu-Wei Zhan, Xin Wang, Hong Chen +6

Video Large Language Models (Video LLMs) have shown impressive performance across a wide range of video-language tasks. However, they often fail in scenarios requiring a deeper und…

cs.RO2025

EvolvingAgent: Curriculum Self-evolving Agent with Continual World Model for Long-Horizon Tasks

Tongtong Feng, Xin Wang, Zekai Zhou +5

Completing Long-Horizon (LH) tasks in open-ended worlds is an important yet difficult problem for embodied agents. Existing approaches suffer from two key challenges: (1) they heav…

cs.CV202422 cited

Multi-weather Cross-view Geo-localization Using Denoising Diffusion Models

Tongtong Feng, Qing Li, Xin Wang +3

Cross-view geo-localization in GNSS-denied environments aims to determine an unknown location by matching drone-view images with the correct geo-tagged satellite-view images from a…