4 papers
VLA-RAIL: A Real-Time Asynchronous Inference Linker for VLA Models and Robots
Yongsheng Zhao, Lei Zhao, Baoping Cheng +3
Vision-Language-Action (VLA) models have achieved remarkable breakthroughs in robotics, with the action chunk playing a dominant role in these advances. Given the real-time and con…
Step-GUI Technical Report
Haolong Yan, Jia Wang, Xin Huang +95
Recent advances in multimodal large language models unlock unprecedented opportunities for GUI automation. However, a fundamental challenge remains: how to efficiently acquire high…
Step-Audio 2 Technical Report
Boyong Wu, Chao Yan, Chen Hu +106
This paper presents Step-Audio 2, an end-to-end multi-modal large language model designed for industry-strength audio understanding and speech conversation. By integrating a latent…
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
Ailin Huang, Boyong Wu, Bruce Wang +142
Real-time speech interaction, serving as a fundamental interface for human-machine collaboration, holds immense potential. However, current open-source models face limitations such…