1 citations · 1 across the 2 of their papers we have counts for
5 papers
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
Shuailei Ma, Jiaqi Liao, Xinyang Wang +24
Despite the recent promise in robot control, video generative models suffer from a domain mismatch due to their primary focus on content creation. For example, their design inheren…
From Foundation to Application: Improving VLA Models in Practice
Wei Wu, Fangjing Wang, Fan Lu +21
Despite recent progress of VLA foundation models, the disparity between laboratory conditions and real-world applications continues to impede their practical implementation. To bri…
Vision Pretraining for Dense Spatial Perception
Zelin Fu, Bin Tan, Changjiang Sun +6
Dense spatial perception is essential for physical intelligence, where visual systems are expected to recover structured, metric, and actionable representations from pixel observat…
Learning Consistent Taxonomic Classification through Hierarchical Reasoning
Zhenghong Li, Kecheng Zheng, Haibin Ling
While Vision-Language Models (VLMs) excel at visual understanding, they often fail to grasp hierarchical knowledge. This leads to common errors where VLMs misclassify coarser taxon…
MePT: Multi-Representation Guided Prompt Tuning for Vision-Language Model
Xinyang Wang, Yi Yang, Minfeng Zhu +3
Recent advancements in pre-trained Vision-Language Models (VLMs) have highlighted the significant potential of prompt tuning for adapting these models to a wide range of downstream…