16 papers
Scene-SAM3D: Multi-View Scene Asset Generation Without Fine-Tuning
Yuqi Zhang, Yadan Luo, Xiangyu Sun +3
High-quality 3D scene assets are critical for embodied applications such as robotic manipulation, navigation, and simulation. Despite their strong object priors, recent single-imag…
Geometry-Guided Self-Supervision for Ultra-Fine-Grained Recognition with Limited Data
Shijie Wang, Yadan Luo, Zijian Wang +3
This paper investigates the intrinsic geometrical features of highly similar objects and introduces a general self-supervised framework called the Geometric Attribute Exploration N…
From Simulation to the Real-World: An In-Field 6D Pose Dataset and Baseline for Robotic Strawberry Harvesting
Woojung Son, Won Suk Lee, Zijing Huang +4
Robotic strawberry harvesting requires precise 6D pose estimation; however, collecting 6D pose ground truth in real agricultural fields is inherently challenging. Existing strawber…
Divide-and-Conquer Approach to Holistic Cognition in High-Similarity Contexts with Limited Data
Shijie Wang, Zijian Wang, Yadan Luo +3
Ultra-fine-grained visual categorization (Ultra-FGVC) aims to classify highly similar subcategories within fine-grained objects using limited training samples. However, holistic ye…
AnchorVLA: Anchored Diffusion for Efficient End-to-End Mobile Manipulation
Jia Syuen Lim, Zhizhen Zhang, Peter Bohm +3
A central challenge in mobile manipulation is preserving multiple plausible action models while remaining reactive during execution. A bottle in a cluttered scene can often be appr…
TALO: Pushing 3D Vision Foundation Models Towards Globally Consistent Online Reconstruction
Fengyi Zhang, Tianjun Zhang, Kasra Khosoussi +3
3D vision foundation models have shown strong generalization in reconstructing key 3D attributes from uncalibrated images through a single feed-forward pass. However, when deployed…