4 papers
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation
Jiasong Xiao, Yutao She, Kai Li +2
Vision-language-action (VLA) models integrate visual observations and language instructions to predict robot actions, demonstrating promising generalization in manipulation tasks.…
MDD-Thinker: Towards Large Reasoning Models for Major Depressive Disorder Diagnosis
Yuyang Sha, Hongxin Pan, Gang Luo +3
Background Major depressive disorder (MDD) is a leading cause of global disability, yet current diagnostic approaches often rely on subjective assessments and lack the ability to i…
MDD-LLM: Towards Accuracy Large Language Models for Major Depressive Disorder Diagnosis
Yuyang Sha, Hongxin Pan, Wei Xu +7
Major depressive disorder (MDD) impacts more than 300 million people worldwide, highlighting a significant public health issue. However, the uneven distribution of medical resource…
Multimodal Perception System for Real Open Environment
Yuyang Sha
This paper presents a novel multimodal perception system for a real open environment. The proposed system includes an embedded computation platform, cameras, ultrasonic sensors, GP…