3 papers
cs.CV2025
Describe Anything Model for Visual Question Answering on Text-rich Images
Yen-Linh Vu, Dinh-Thang Duong, Truong-Binh Duong +8
Recent progress has been made in region-aware vision-language modeling, particularly with the emergence of the Describe Anything Model (DAM). DAM is capable of generating detailed…
cs.CL2025
ViQA-COVID: COVID-19 Machine Reading Comprehension Dataset for Vietnamese
Hai-Chung Nguyen-Phung, Ngoc C. Lê, Van-Chien Nguyen +2
After two years of appearance, COVID-19 has negatively affected people and normal life around the world. As in May 2022, there are more than 522 million cases and six million death…
cs.CV2025
Leveraging Habitat Information for Fine-grained Bird Identification
Tin Nguyen, Peijie Chen, Anh Totti Nguyen
Traditional bird classifiers mostly rely on the visual characteristics of birds. Some prior works even train classifiers to be invariant to the background, completely discarding th…