11 papers
RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination
Haotian Liang, Mingkang Chen, Yufei Huang +27
The paper introduces RxBrain, a foundation model that jointly reasons over language and visual inputs to create embodied plans, using a multimodal Mixture-of-Transformers architect…
Hy-Embodied-VLM-1.0: Efficient Physical-World Agents
Ziyi Wang, Xumin Yu, Yongming Rao +19
Building capable embodied agents requires not only multimodal perception and understanding, but also agentic capabilities for reasoning about actions, adapting to evolving situatio…
Hy-Embodied-0.5-VLA: From Vision-Language-Action Models to a Real-World Robot Learning Stack
He Zhang, Lingzhu Xiang, Haitao Lin +23
In this report, we present Hy-Embodied-0.5-VLA, abbreviated as HyVLA-0.5, an end-to-end system that spans the full robot learning stack: data collection, model design, continued pr…
GEM: Generative Supervision Helps Embodied Intelligence
Ruowen Zhao, Bangguo Li, Zuyan Liu +9
Embodied Vision-Language Models (VLMs) have demonstrated impressive performance and generalization in robotics, particularly within Vision-Language-Action frameworks. However, a si…
Towards Real-World Document Parsing via Realistic Scene Synthesis and Document-Aware Training
Gengluo Li, Pengyuan Lyu, Chengquan Zhang +7
Document parsing has recently advanced with multimodal large language models (MLLMs) that directly map document images to structured outputs. Traditional cascaded pipelines depend…
Intern-S1-Pro: Scientific Multimodal Foundation Model at Trillion Scale
Yicheng Zou, Dongsheng Zhu, Lin Zhu +174
We introduce Intern-S1-Pro, the first one-trillion-parameter scientific multimodal foundation model. Scaling to this unprecedented size, the model delivers a comprehensive enhancem…