14 citations · 20 across the 4 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2024★ 5 cited
SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
Boyuan Chen, Zhuo Xu, Sean Kirmani +6
Understanding and reasoning about spatial relationships is a fundamental capability for Visual Question Answering (VQA) and robotics. While Vision Language Models (VLM) have demons…
cs.CV2023★ 14 cited
Understanding Why ViT Trains Badly on Small Datasets: An Intuitive Perspective
Haoran Zhu, Boyuan Chen, Carter Yang
Vision transformer (ViT) is an attention neural network architecture that is shown to be effective for computer vision tasks. However, compared to ResNet-18 with a similar number o…