1 citations · 1 across the 6 of their papers we have counts for
4 papers · 1 filter
Look Before You Zoom: Adaptive Routing for the Resolution-Context Trade-off in Visual RAG
Oanh N. Tran, Thanh Quoc Hung Le, Oscar Chew +2
Vision-Language Models (VLMs) struggle as query-relevant objects become smaller. To address this, recent training-free approaches dynamically retrieve and zoom into local image reg…
Pop-Up Distractions Reveal Bag-of-Events Behavior in Video Large Language Models
Oscar Chew, Serhii Honcharenko, Qian-Hui Chen +4
A key capability for video understanding is reliably linking subjects to events across time, yet whether Video Large Language Models (VideoLLMs) actually achieve this remains uncle…
Is CLIP Cross-Eyed? Revealing and Mitigating Center Bias in the CLIP Family
Oscar Chew, Hsiao-Ying Huang, Kunal Jain +3
Recent research has shown that contrastive vision-language models such as CLIP often lack fine-grained understanding of visual content. While a growing body of work has sought to a…
Defending Text-to-image Diffusion Models: Surprising Efficacy of Textual Perturbations Against Backdoor Attacks
Oscar Chew, Po-Yi Lu, Jayden Lin +1
Text-to-image diffusion models have been widely adopted in real-world applications due to their ability to generate realistic images from textual descriptions. However, recent stud…