2 papers
cs.CV2025
GaussianToken: An Effective Image Tokenizer with 2D Gaussian Splatting
Jiajun Dong, Chengkun Wang, Wenzhao Zheng +3
Effective image tokenization is crucial for both multi-modal understanding and generation tasks due to the necessity of the alignment with discrete text data. To this end, existing…
cs.CV2025
Localization-Aware Multi-Scale Representation Learning for Repetitive Action Counting
Sujia Wang, Xiangwei Shen, Yansong Tang +3
Repetitive action counting (RAC) aims to estimate the number of class-agnostic action occurrences in a video without exemplars. Most current RAC methods rely on a raw frame-to-fram…