activity
20232026
most citedA Cosine Similarity-based Method for Out-of-Distribution Detection

2 citations · 8 across the 37 of their papers we have counts for

collaborators
Showing cs.CVShow all

10 papers · 1 filter

cs.CV2026

Look Before You Zoom: Adaptive Routing for the Resolution-Context Trade-off in Visual RAG

Oanh N. Tran, Thanh Quoc Hung Le, Oscar Chew +2

Vision-Language Models (VLMs) struggle as query-relevant objects become smaller. To address this, recent training-free approaches dynamically retrieve and zoom into local image reg…

cs.CV2026

FOCA: Future-Oriented Conditioning for Data-Efficient Vision-Language-Action Adaptation

Duc Minh Nguyen, Nghiem Tuong Diep, Binh Gia Nguyen +20

Vision-Language-Action (VLA) models enable general-purpose robotic control via large-scale multimodal pretraining, yet their effectiveness under few-shot imitation learning remains…

cs.CV2026

SparseSAM: Structured Sparsification of Activations in Segment Anything Models

Hoai-Chau Tran, Chi H. Nguyen, Duy M. H. Nguyen +3

The Segment Anything Model (SAM) achieves strong open-vocabulary segmentation, but its ViT-based image encoders dominate inference latency and memory. Existing activation compressi…

cs.CV2026

Pop-Up Distractions Reveal Bag-of-Events Behavior in Video Large Language Models

Oscar Chew, Serhii Honcharenko, Qian-Hui Chen +4

A key capability for video understanding is reliably linking subjects to events across time, yet whether Video Large Language Models (VideoLLMs) actually achieve this remains uncle…

cs.CV2026

Is CLIP Cross-Eyed? Revealing and Mitigating Center Bias in the CLIP Family

Oscar Chew, Hsiao-Ying Huang, Kunal Jain +3

Recent research has shown that contrastive vision-language models such as CLIP often lack fine-grained understanding of visual content. While a growing body of work has sought to a…

cs.CV2026

StructSAM: Structure- and Spectrum-Preserving Token Merging for Segment Anything Models

Duy M. H. Nguyen, Tuan A. Tran, Duong Nguyen +17

Recent token merging techniques for Vision Transformers (ViTs) provide substantial speedups by reducing the number of tokens processed by self-attention, often without retraining.…