3 papers
cs.RO2025
Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions
Wei Zhao, Gongsheng Li, Zhefei Gong +3
Vision-Language-Action (VLA) models have recently become highly prominent in the field of robotics. Leveraging vision-language foundation models trained on large-scale internet dat…
cs.CV2024
StyleTex: Style Image-Guided Texture Generation for 3D Models
Zhiyu Xie, Yuqing Zhang, Xiangjun Tang +4
Style-guided texture generation aims to generate a texture that is harmonious with both the style of the reference image and the geometry of the input mesh, given a reference style…
cs.CV2024
Promoting CNNs with Cross-Architecture Knowledge Distillation for Efficient Monocular Depth Estimation
Zhimeng Zheng, Tao Huang, Gongsheng Li +1
Recently, the performance of monocular depth estimation (MDE) has been significantly boosted with the integration of transformer models. However, the transformer models are usually…