most citedMePT: Multi-Representation Guided Prompt Tuning for Vision-Language Model

1 citations · 1 across the 2 of their papers we have counts for

collaborators

5 papers

cs.CV2026

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence

Shuailei Ma, Jiaqi Liao, Xinyang Wang +24

Despite the recent promise in robot control, video generative models suffer from a domain mismatch due to their primary focus on content creation. For example, their design inheren…

cs.RO2026

From Foundation to Application: Improving VLA Models in Practice

Wei Wu, Fangjing Wang, Fan Lu +21

Despite recent progress of VLA foundation models, the disparity between laboratory conditions and real-world applications continues to impede their practical implementation. To bri…

cs.CV2026

Vision Pretraining for Dense Spatial Perception

Zelin Fu, Bin Tan, Changjiang Sun +6

Dense spatial perception is essential for physical intelligence, where visual systems are expected to recover structured, metric, and actionable representations from pixel observat…

cs.CV2026

Learning Consistent Taxonomic Classification through Hierarchical Reasoning

Zhenghong Li, Kecheng Zheng, Haibin Ling

While Vision-Language Models (VLMs) excel at visual understanding, they often fail to grasp hierarchical knowledge. This leads to common errors where VLMs misclassify coarser taxon…

cs.CV20241 cited

MePT: Multi-Representation Guided Prompt Tuning for Vision-Language Model

Xinyang Wang, Yi Yang, Minfeng Zhu +3

Recent advancements in pre-trained Vision-Language Models (VLMs) have highlighted the significant potential of prompt tuning for adapting these models to a wide range of downstream…