2 citations · 2 across the 7 of their papers we have counts for
1 paper · 1 filter
Jianyu Zhang, Li Zhang, Shijian Li
The visual understanding are often approached from 3 granular levels: image, patch and pixel. Visual Tokenization, trained by self-supervised reconstructive learning, compresses vi…