2 citations · 2 across the 3 of their papers we have counts for
3 papers
cs.RO2025
Evo-1: Lightweight Vision-Language-Action Model with Preserved Semantic Alignment
Tao Lin, Yilei Zhong, Yuxin Du +11
Vision-Language-Action (VLA) models have emerged as a powerful framework that unifies perception, language, and control, enabling robots to perform diverse tasks through multimodal…
cs.MM2025
Multi-modal and Metadata Capture Model for Micro Video Popularity Prediction
Jiacheng Lu, Mingyuan Xiao, Weijian Wang +3
As short videos have become the primary form of content consumption across various industries, accurately predicting their popularity has become key to enhancing user engagement an…
cs.CV2024★ 2 cited
Touchstone Benchmark: Are We on the Right Way for Evaluating AI Algorithms for Medical Segmentation?
Pedro R. A. S. Bassi, Wenxuan Li, Yucheng Tang +50
How can we test AI performance? This question seems trivial, but it isn't. Standard benchmarks often have problems such as in-distribution and small-size test sets, oversimplified…