3 citations · 6 across the 3 of their papers we have counts for
3 papers · 1 filter
What You See is What You Read? Improving Text-Image Alignment Evaluation
Michal Yarom, Yonatan Bitton, Soravit Changpinyo +5
Automatically determining whether a text and a corresponding image are semantically aligned is a significant challenge for vision-language models, with applications in generative t…
Decoupled Box Proposal and Featurization with Ultrafine-Grained Semantic Labels Improve Image Captioning and Visual Question Answering
Soravit Changpinyo, Bo Pang, Piyush Sharma +1
Object detection plays an important role in current solutions to vision and language tasks like image captioning and visual question answering. However, popular models like Faster…
Multi-Task Learning for Sequence Tagging: An Empirical Study
Soravit Changpinyo, Hexiang Hu, Fei Sha
We study three general multi-task learning (MTL) approaches on 11 sequence tagging tasks. Our extensive empirical results show that in about 50% of the cases, jointly learning all…