10 papers
From Foundation to Application: Improving VLA Models in Practice
Wei Wu, Fangjing Wang, Fan Lu +21
Despite recent progress of VLA foundation models, the disparity between laboratory conditions and real-world applications continues to impede their practical implementation. To bri…
Vision Pretraining for Dense Spatial Perception
Zelin Fu, Bin Tan, Changjiang Sun +6
Dense spatial perception is essential for physical intelligence, where visual systems are expected to recover structured, metric, and actionable representations from pixel observat…
Sat3DGen: Comprehensive Street-Level 3D Scene Generation from Single Satellite Image
Ming Qian, Zimin Xia, Changkun Liu +6
Generating a street-level 3D scene from a single satellite image is a crucial yet challenging task. Current methods present a stark trade-off: geometry-colorization models achieve…
DepthLab: From Partial to Complete
Zhiheng Liu, Ka Leong Cheng, Qiuyu Wang +7
Missing values remain a common challenge for depth data across its wide range of applications, stemming from various causes like incomplete data acquisition and perspective alterat…
Seeing through Satellite Images at Street Views
Ming Qian, Bin Tan, Qiuyu Wang +5
This paper studies the task of SatStreet-view synthesis, which aims to render photorealistic street-view panorama images and videos given any satellite image and specified camera p…
Interacted Planes Reveal 3D Line Mapping
Zeran Ke, Bin Tan, Gui-Song Xia +2
3D line mapping from multi-view RGB images provides a compact and structured visual representation of scenes. We study the problem from a physical and topological perspective: a 3D…