3 citations · 4 across the 3 of their papers we have counts for
5 papers · 1 filter
Understanding Segment Anything Model: SAM is Biased Towards Texture Rather than Shape
Chaoning Zhang, Yu Qiao, Shehbaz Tariq +5
In contrast to the human vision that mainly depends on the shape for recognizing the objects, deep image recognition models are widely known to be biased toward texture. Recently,…
Toward a Deeper Understanding: RetNet Viewed through Convolution
Chenghao Li, Chaoning Zhang
The success of Vision Transformer (ViT) has been widely reported on a wide range of image recognition tasks. ViT can learn global dependencies superior to CNN, yet CNN's inherent l…
A Survey on Segment Anything Model (SAM): Vision Foundation Model Meets Prompt Engineering
Chaoning Zhang, Joseph Cho, Fachrina Dewi Puspitasari +11
The Segment Anything Model (SAM), developed by Meta AI Research, represents a significant breakthrough in computer vision, offering a robust framework for image and video segmentat…
When ChatGPT for Computer Vision Will Come? From 2D to 3D
Chenghao Li, Chaoning Zhang
ChatGPT and its improved variant GPT4 have revolutionized the NLP field with a single model solving almost all text related tasks. However, such a model for computer vision does no…
Generative AI meets 3D: A Survey on Text-to-3D in AIGC Era
Chenghao Li, Chaoning Zhang, Joseph Cho +6
Generative AI has made significant progress in recent years, with text-guided content generation being the most practical as it facilitates interaction between human instructions a…