3 papers
cs.CV2026
Scene-SAM3D: Multi-View Scene Asset Generation Without Fine-Tuning
Yuqi Zhang, Yadan Luo, Xiangyu Sun +3
High-quality 3D scene assets are critical for embodied applications such as robotic manipulation, navigation, and simulation. Despite their strong object priors, recent single-imag…
cs.CV2025
DualCamCtrl: Dual-Branch Diffusion Model for Geometry-Aware Camera-Controlled Video Generation
Hongfei Zhang, Kanghao Chen, Zixin Zhang +6
This paper presents DualCamCtrl, a novel end-to-end diffusion model for camera-controlled video generation. Recent works have advanced this field by representing camera poses as ra…
cs.RO2025
MergeVLA: Cross-Skill Model Merging Toward a Generalist Vision-Language-Action Agent
Yuxia Fu, Zhizhen Zhang, Yuqi Zhang +3
Recent Vision-Language-Action (VLA) models reformulate vision-language models by tuning them with millions of robotic demonstrations. While they perform well when fine-tuned for a…