13 papers
CompassNav: Steering From Path Imitation To Decision Understanding In Navigation
LinFeng Li, Jian Zhao, Yuan Xie +2
The dominant paradigm for training Large Vision-Language Models (LVLMs) in navigation relies on imitating expert trajectories. This approach reduces the complex navigation task to…
Visual Attention Reasoning via Hierarchical Search and Self-Verification
Wei Cai, Jian Zhao, Yuchen Yuan +4
Multimodal Large Language Models (MLLMs) frequently hallucinate due to their reliance on fragile, linear reasoning and weak visual grounding. We propose Visual Attention Reasoning…
Autonomous Driving in Unstructured Environments: How Far Have We Come?
Chen Min, Shubin Si, Xu Wang +14
Research on autonomous driving in unstructured outdoor environments is less advanced than in structured urban settings due to challenges like environmental diversities and scene co…
Loupe: A Generalizable and Adaptive Framework for Image Forgery Detection
Yuchu Jiang, Jiaming Chu, Jian Zhao +5
The proliferation of generative models has raised serious concerns about visual content forgery. Existing deepfake detection methods primarily target either image-level classificat…
Aetheria: A multimodal interpretable content safety framework based on multi-agent debate and collaboration
Yuxiang He, Jian Zhao, Yuchen Yuan +8
The exponential growth of digital content presents significant challenges for content safety. Current moderation systems, often based on single models or fixed pipelines, exhibit l…
A Parameter-Efficient Mixture-of-Experts Framework for Cross-Modal Geo-Localization
LinFeng Li, Jian Zhao, Zepeng Yang +6
We present a winning solution to RoboSense 2025 Track 4: Cross-Modal Drone Navigation. The task retrieves the most relevant geo-referenced image from a large multi-platform corpus…