1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.AI2024★ 1 cited
VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents
Xiao Liu, Tianjie Zhang, Yu Gu +27
Large Multimodal Models (LMMs) have ushered in a new era in artificial intelligence, merging capabilities in both language and vision to form highly capable Visual Foundation Agent…
cs.RO2023
STOPNet: Multiview-based 6-DoF Suction Detection for Transparent Objects on Production Lines
Yuxuan Kuang, Qin Han, Danshi Li +5
In this work, we present STOPNet, a framework for 6-DoF object suction detection on production lines, with a focus on but not limited to transparent objects, which is an important…