5 papers · 1 filter
Spatially-Aware Class-Agnostic Object Counting
Robert Wijaya, Md. Tanvir Hossain, Amanda Kau +1
Generalised object counting aims to estimate the number of instances of an arbitrary object category from a single image, but many recent methods can struggle on structurally compl…
Mixture of Cognitive Experts in Large Vision-Language Models
Robert Wijaya, Ngai-Man Cheung
Large Vision Language Models (LVLMs) require strong reasoning over both visual and textual input. Recent work suggests that cognitive elements, especially diverse representations a…
Crowdsource, Crawl, or Generate? Creating SEA-VL, a Multicultural Vision-Language Dataset for Southeast Asia
Samuel Cahyawijaya, Holy Lovenia, Joel Ruben Antony Moniz +89
Southeast Asia (SEA) is a region of extraordinary linguistic and cultural diversity, yet it remains significantly underrepresented in vision-language (VL) research. This often resu…
Synth-Align: Improving Trustworthiness in Vision-Language Model with Synthetic Preference Data Alignment
Robert Wijaya, Ngoc-Bao Nguyen, Ngai-Man Cheung
Large Vision-Language Models (LVLMs) have shown promising capabilities in understanding and generating information by integrating both visual and textual data. However, current mod…
Investigating the Robustness and Properties of Detection Transformers (DETR) Toward Difficult Images
Zhao Ning Zou, Yuhang Zhang, Robert Wijaya
Transformer-based object detectors (DETR) have shown significant performance across machine vision tasks, ultimately in object detection. This detector is based on a self-attention…