4 papers
Describe Anything Model for Visual Question Answering on Text-rich Images
Yen-Linh Vu, Dinh-Thang Duong, Truong-Binh Duong +8
Recent progress has been made in region-aware vision-language modeling, particularly with the emergence of the Describe Anything Model (DAM). DAM is capable of generating detailed…
ViQA-COVID: COVID-19 Machine Reading Comprehension Dataset for Vietnamese
Hai-Chung Nguyen-Phung, Ngoc C. Lê, Van-Chien Nguyen +2
After two years of appearance, COVID-19 has negatively affected people and normal life around the world. As in May 2022, there are more than 522 million cases and six million death…
PEEB: Part-based Image Classifiers with an Explainable and Editable Language Bottleneck
Thang M. Pham, Peijie Chen, Tin Nguyen +3
CLIP-based classifiers rely on the prompt containing a {class name} that is known to the text encoder. Therefore, they perform poorly on new classes or the classes whose names rare…
Leveraging Habitat Information for Fine-grained Bird Identification
Tin Nguyen, Peijie Chen, Anh Totti Nguyen
Traditional bird classifiers mostly rely on the visual characteristics of birds. Some prior works even train classifiers to be invariant to the background, completely discarding th…