4 papers · 1 filter
City-VLM: Towards Multidomain Perception Scene Understanding via Multimodal Incomplete Learning
Penglei Sun, Yaoxian Song, Xiangru Zhu +7
Scene understanding enables intelligent agents to interpret and comprehend their environment. While existing large vision-language models (LVLMs) for scene understanding have prima…
OG-Mapping: Octree-based Structured 3D Gaussians for Online Dense Mapping
Meng Wang, Junyi Wang, Changqun Xia +2
3D Gaussian splatting (3DGS) has recently demonstrated promising advancements in RGB-D online dense mapping. Nevertheless, existing methods excessively rely on per-pixel depth cues…
PGNeXt: High-Resolution Salient Object Detection via Pyramid Grafting Network
Changqun Xia, Chenxi Xie, Zhentao He +2
We present an advanced study on more challenging high-resolution salient object detection (HRSOD) from both dataset and network framework perspectives. To compensate for the lack o…
Towards Imbalanced Motion: Part-Decoupling Network for Video Portrait Segmentation
Tianshu Yu, Changqun Xia, Jia Li
Video portrait segmentation (VPS), aiming at segmenting prominent foreground portraits from video frames, has received much attention in recent years. However, simplicity of existi…