7 papers
Hy-Embodied-VLM-1.0: Efficient Physical-World Agents
Ziyi Wang, Xumin Yu, Yongming Rao +19
Building capable embodied agents requires not only multimodal perception and understanding, but also agentic capabilities for reasoning about actions, adapting to evolving situatio…
ViQ: Text-Aligned Visual Quantized Representations at Any Resolution
Xumin Yu, Zuyan Liu, Zhenyu Yang +5
A unified representation for text and vision is a natural pursuit, as it enables simpler multimodal modeling and more efficient training. However, representing images as discrete s…
Hy-Embodied-0.5-VLA: From Vision-Language-Action Models to a Real-World Robot Learning Stack
He Zhang, Lingzhu Xiang, Haitao Lin +23
In this report, we present Hy-Embodied-0.5-VLA, abbreviated as HyVLA-0.5, an end-to-end system that spans the full robot learning stack: data collection, model design, continued pr…
GEM: Generative Supervision Helps Embodied Intelligence
Ruowen Zhao, Bangguo Li, Zuyan Liu +9
Embodied Vision-Language Models (VLMs) have demonstrated impressive performance and generalization in robotics, particularly within Vision-Language-Action frameworks. However, a si…
HY-Embodied-0.5: Embodied Foundation Models for Real-World Agents
Tencent Robotics X, HY Vision Team, : +20
We introduce HY-Embodied-0.5, a family of foundation models specifically designed for real-world embodied agents. To bridge the gap between general Vision-Language Models (VLMs) an…
Spatial-TTT: Streaming Visual-based Spatial Intelligence with Test-Time Training
Fangfu Liu, Diankun Wu, Jiawei Chi +7
Humans perceive and understand real-world spaces through a stream of visual observations. Therefore, the ability to streamingly maintain and update spatial evidence from potentiall…