6 papers
StructSAM: Structure- and Spectrum-Preserving Token Merging for Segment Anything Models
Duy M. H. Nguyen, Tuan A. Tran, Duong Nguyen +17
Recent token merging techniques for Vision Transformers (ViTs) provide substantial speedups by reducing the number of tokens processed by self-attention, often without retraining.…
FOCA: Future-Oriented Conditioning for Data-Efficient Vision-Language-Action Adaptation
Duc Minh Nguyen, Nghiem Tuong Diep, Binh Gia Nguyen +20
Vision-Language-Action (VLA) models enable general-purpose robotic control via large-scale multimodal pretraining, yet their effectiveness under few-shot imitation learning remains…
A Hybrid Vision Transformer Approach for Mathematical Expression Recognition
Anh Duy Le, Van Linh Pham, Vinh Loi Ly +3
One of the crucial challenges taken in document analysis is mathematical expression recognition. Unlike text recognition which only focuses on one-dimensional structure images, mat…
CONSTANT: Towards High-Quality One-Shot Handwriting Generation with Patch Contrastive Enhancement and Style-Aware Quantization
Anh-Duy Le, Van-Linh Pham, Thanh-Nam Vo +2
One-shot styled handwriting image generation, despite achieving impressive results in recent years, remains challenging due to the difficulty in capturing the intricate and diverse…
How Many Tokens Do 3D Point Cloud Transformer Architectures Really Need?
Tuan Anh Tran, Duy M. H. Nguyen, Hoai-Chau Tran +7
Recent advances in 3D point cloud transformers have led to state-of-the-art results in tasks such as semantic segmentation and reconstruction. However, these models typically rely…
SepFormer: Coarse-to-fine Separator Regression Network for Table Structure Recognition
Nam Quan Nguyen, Xuan Phong Pham, Tuan-Anh Tran
The automated reconstruction of the logical arrangement of tables from image data, termed Table Structure Recognition (TSR), is fundamental for semantic data extraction. Recently,…