16 papers
StructSAM: Structure- and Spectrum-Preserving Token Merging for Segment Anything Models
Duy M. H. Nguyen, Tuan A. Tran, Duong Nguyen +17
Recent token merging techniques for Vision Transformers (ViTs) provide substantial speedups by reducing the number of tokens processed by self-attention, often without retraining.…
Self-Improving VLA Policies: Selected Diffusion Noise for Spurious-Robust Action Smoothing
Duc Minh Nguyen, Bao-Ngoc Dao, Tung M. Luu +15
Diffusion-based Vision-Language-Action (VLA) policies enable strong generalization in robotic manipulation, but remain sensitive to spurious visual correlations and noisy action ge…
Protein Fold Classification at Scale: Benchmarking and Pretraining
Dexiong Chen, Andrei Manolache, Mathias Niepert +1
Classifying protein topology is essential for deciphering biological function, but progress is held back by the lack of large-scale benchmarks that avoid duplicates and by models t…
SparseSAM: Structured Sparsification of Activations in Segment Anything Models
Hoai-Chau Tran, Chi H. Nguyen, Duy M. H. Nguyen +3
The Segment Anything Model (SAM) achieves strong open-vocabulary segmentation, but its ViT-based image encoders dominate inference latency and memory. Existing activation compressi…
Logical Guidance for the Exact Composition of Diffusion Models
Francesco Alesiani, Jonathan Warrell, Tanja Bien +3
We propose LOGDIFF (Logical Guidance for the Exact Composition of Diffusion Models), a guidance framework for diffusion models that enables principled constrained generation with c…
Adaptive Width Neural Networks
Federico Errica, Henrik Christiansen, Viktor Zaverkin +2
For almost 70 years, researchers have typically selected the width of neural networks' layers either manually or through automated hyperparameter tuning methods such as grid search…