3 papers
cs.CV2025
GaussianToken: An Effective Image Tokenizer with 2D Gaussian Splatting
Jiajun Dong, Chengkun Wang, Wenzhao Zheng +3
Effective image tokenization is crucial for both multi-modal understanding and generation tasks due to the necessity of the alignment with discrete text data. To this end, existing…
cs.CV2024
V2M: Visual 2-Dimensional Mamba for Image Representation Learning
Chengkun Wang, Wenzhao Zheng, Yuanhui Huang +2
Mamba has garnered widespread attention due to its flexible design and efficient hardware performance to process 1D sequences based on the state space model (SSM). Recent studies h…
cs.CV2024
GlobalMamba: Global Image Serialization for Vision Mamba
Chengkun Wang, Wenzhao Zheng, Jie Zhou +1
Vision mambas have demonstrated strong performance with linear complexity to the number of vision tokens. Their efficiency results from processing image tokens sequentially. Howeve…