1 citations · 2 across the 11 of their papers we have counts for
4 papers · 1 filter
EgoTac: In-the-wild Tactile Prediction from Egocentric Vision
Wenkang Zhang, Chengbo Yuan, Zicheng Zhang +2
Touch is fundamental to dexterous manipulation, yet most egocentric human data increasingly used for robot learning lacks tactile information. Directly collecting large-scale tacti…
Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes
Zhiyuan Feng, Zhaolu Kang, Qijie Wang +16
Vision-language models (VLMs) are essential to Embodied AI, enabling robots to perceive, reason, and act in complex environments. They also serve as the foundation for the recent V…
Seed1.5-VL Technical Report
Dong Guo, Faming Wu, Feida Zhu +194
We present Seed1.5-VL, a vision-language foundation model designed to advance general-purpose multimodal understanding and reasoning. Seed1.5-VL is composed with a 532M-parameter v…
Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos
Chengbo Yuan, Geng Chen, Li Yi +1
Egocentric videos provide valuable insights into human interactions with the physical world, which has sparked growing interest in the computer vision and robotics communities. A c…