287 citations · 344 across the 20 of their papers we have counts for
33 papers · 1 filter
SphereSOD: Geometry-Structure Coupled Learning for 360 Salient Object Detection
Junsong Zhang, Zhijie Shen, Shuai Zheng +4
360° salient object detection (SOD) aims to accurately segment salient regions across a full field of view. However, equirectangular projection (ERP) introduces severe spatial dist…
Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States
Kang Liao, Yihang Luo, Xiao-Ming Wu +7
We propose Puffin-World, a unified multimodal architecture that integrates physical understanding, spatial simulation, and 3D world generation and reconstruction without relying on…
CaC: Advancing Video Reward Models via Hierarchical Spatiotemporal Concentrating
Jiyuan Wang, Huan Ouyang, Jiuzhou Lin +15
In this paper, we propose Concentrate and Concentrate (CaC), a coarse-to-fine anomaly reward model based on Vision-Language Models. During inference, it first conducts a global tem…
UniStitch: Unifying Semantic and Geometric Features for Image Stitching
Yuan Mei, Lang Nie, Kang Liao +3
Traditional image stitching methods estimate warps from hand-crafted geometric features, whereas recent learning-based solutions leverage semantic features from neural networks ins…
RePer-360: Releasing Perspective Priors for 360 Depth Estimation via Self-Modulation
Cheng Guan, Chunyu Lin, Zhijie Shen +2
Recent depth foundation models trained on perspective imagery achieve strong performance, yet generalize poorly to 360 images due to the substantial geometric discrepancy b…
Edit in 2D, Verify in 3D: Reinforcement Learning for Multi-view Consistent Scene Editing
Jiyuan Wang, Chunyu Lin, Lei Sun +8
Leveraging the priors of 2D diffusion models for 3D editing has emerged as a promising paradigm. However, multi-view consistency remains challenging in edited results, and the extr…