1 citations · 1 across the 2 of their papers we have counts for
3 papers
cs.CY2026
How Many Visual Levers Drive Urban Perception? Interventional Counterfactuals via Multiple Localised Edits
Jason Tang, Stephen Law
Street-view perception models predict subjective attributes such as safety at scale, but remain correlational: they do not identify which localized visual changes would plausibly s…
cs.IR2024
Shopping Queries Image Dataset (SQID): An Image-Enriched ESCI Dataset for Exploring Multimodal Learning in Product Search
Marie Al Ghossein, Ching-Wei Chen, Jason Tang
Recent advances in the fields of Information Retrieval and Machine Learning have focused on improving the performance of search engines to enhance the user experience, especially i…
cs.IR2024★ 1 cited
Captions Are Worth a Thousand Words: Enhancing Product Retrieval with Pretrained Image-to-Text Models
Jason Tang, Garrin McGoldrick, Marie Al-Ghossein +1
This paper explores the usage of multimodal image-to-text models to enhance text-based item retrieval. We propose utilizing pre-trained image captioning and tagging models, such as…