2 citations · 8 across the 37 of their papers we have counts for
10 papers · 1 filter
Look Before You Zoom: Adaptive Routing for the Resolution-Context Trade-off in Visual RAG
Oanh N. Tran, Thanh Quoc Hung Le, Oscar Chew +2
Vision-Language Models (VLMs) struggle as query-relevant objects become smaller. To address this, recent training-free approaches dynamically retrieve and zoom into local image reg…
FOCA: Future-Oriented Conditioning for Data-Efficient Vision-Language-Action Adaptation
Duc Minh Nguyen, Nghiem Tuong Diep, Binh Gia Nguyen +20
Vision-Language-Action (VLA) models enable general-purpose robotic control via large-scale multimodal pretraining, yet their effectiveness under few-shot imitation learning remains…
SparseSAM: Structured Sparsification of Activations in Segment Anything Models
Hoai-Chau Tran, Chi H. Nguyen, Duy M. H. Nguyen +3
The Segment Anything Model (SAM) achieves strong open-vocabulary segmentation, but its ViT-based image encoders dominate inference latency and memory. Existing activation compressi…
Pop-Up Distractions Reveal Bag-of-Events Behavior in Video Large Language Models
Oscar Chew, Serhii Honcharenko, Qian-Hui Chen +4
A key capability for video understanding is reliably linking subjects to events across time, yet whether Video Large Language Models (VideoLLMs) actually achieve this remains uncle…
Is CLIP Cross-Eyed? Revealing and Mitigating Center Bias in the CLIP Family
Oscar Chew, Hsiao-Ying Huang, Kunal Jain +3
Recent research has shown that contrastive vision-language models such as CLIP often lack fine-grained understanding of visual content. While a growing body of work has sought to a…
StructSAM: Structure- and Spectrum-Preserving Token Merging for Segment Anything Models
Duy M. H. Nguyen, Tuan A. Tran, Duong Nguyen +17
Recent token merging techniques for Vision Transformers (ViTs) provide substantial speedups by reducing the number of tokens processed by self-attention, often without retraining.…