3 papers
cs.CV2026
MBench: A Comprehensive Benchmark on Memory Capability for Video World Models
Shengjun Zhang, Zhang Zhang, Simin Huang +11
Recent advancements in video-based world models have demonstrated an unprecedented ability to synthesize high-fidelity visual sequences. However, a fundamental gap persists between…
cs.CV2023
HGDNet: A Height-Hierarchy Guided Dual-Decoder Network for Single View Building Extraction and Height Estimation
Chaoran Lu, Ningning Cao, Pan Zhang +7
Unifying the correlative single-view satellite image building extraction and height estimation tasks indicates a promising way to share representations and acquire generalist model…
cs.CV2023
Fine-grained building roof instance segmentation based on domain adapted pretraining and composite dual-backbone
Guozhang Liu, Baochai Peng, Ting Liu +7
The diversity of building architecture styles of global cities situated on various landforms, the degraded optical imagery affected by clouds and shadows, and the significant inter…