From the 1 of 14 linked papers with an AI index.
14 papers
GeoChrono: Benchmarking and Rethinking Long-Term Temporal Understanding in Remote Sensing
Yujie Li, Jiancheng Pan, Zhiwei Wei +3
Remote sensing offers an unparalleled vantage point for observing the Earth's long-term surface evolution, yet it demands that a model not only perceive land cover at isolated mome…
Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence
Zhishan Zou, Guoyan Sun, Zhiwei Wei +4
The paper presents SIS-Bench, a benchmark for assessing UAV embodied spatial intelligence that jointly evaluates spatial cognition and self‑awareness across perception, memory, and…
The Second Challenge on Cross-Domain Few-Shot Object Detection at NTIRE 2026: Methods and Results
Xingyu Qiu, Yuqian Fu, Jiawei Geng +70
Cross-domain few-shot object detection (CD-FSOD) remains a challenging problem for existing object detectors and few-shot learning approaches, particularly when generalizing across…
V-SAM: Marrying SAM2 with Multi-Prompt Experts for Cross-View Object Correspondence
Jiancheng Pan, Runze Wang, Tianwen Qian +7
Cross-view object correspondence, exemplified by the representative task of ego-exo object correspondence, aims to establish consistent associations of the same object across diffe…
Epistemic-aware Vision-Language Foundation Model for Fetal Ultrasound Interpretation
Xiao He, Huangxuan Zhao, Guojia Wan +8
Recent medical vision-language models have shown promise on tasks such as VQA, report generation, and anomaly detection. However, most are adapted to structured adult imaging and u…
Exo2EgoSyn: Unlocking Foundation Video Generation Models for Exocentric-to-Egocentric Video Synthesis
Mohammad Mahdi, Yuqian Fu, Nedko Savov +3
Foundation video generation models such as WAN 2.2 exhibit strong text- and image-conditioned synthesis abilities but remain constrained to the same-view generation setting. In thi…