4 papers
Spatially-Aware Class-Agnostic Object Counting
Robert Wijaya, Md. Tanvir Hossain, Amanda Kau +1
Generalised object counting aims to estimate the number of instances of an arbitrary object category from a single image, but many recent methods can struggle on structurally compl…
Mixture of Cognitive Experts in Large Vision-Language Models
Robert Wijaya, Ngai-Man Cheung
Large Vision Language Models (LVLMs) require strong reasoning over both visual and textual input. Recent work suggests that cognitive elements, especially diverse representations a…
Synth-Align: Improving Trustworthiness in Vision-Language Model with Synthetic Preference Data Alignment
Robert Wijaya, Ngoc-Bao Nguyen, Ngai-Man Cheung
Large Vision-Language Models (LVLMs) have shown promising capabilities in understanding and generating information by integrating both visual and textual data. However, current mod…
Crowdsource, Crawl, or Generate? Creating SEA-VL, a Multicultural Vision-Language Dataset for Southeast Asia
Samuel Cahyawijaya, Holy Lovenia, Joel Ruben Antony Moniz +89
Southeast Asia (SEA) is a region of extraordinary linguistic and cultural diversity, yet it remains significantly underrepresented in vision-language (VL) research. This often resu…