1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.CV2023★ 1 cited
FashionNTM: Multi-turn Fashion Image Retrieval via Cascaded Memory
Anwesan Pal, Sahil Wadhwa, Ayush Jaiswal +5
Multi-turn textual feedback-based fashion image retrieval focuses on a real-world setting, where users can iteratively provide information to refine retrieval results until they fi…
cs.CV2023
MoMo: A shared encoder Model for text, image and multi-Modal representations
Rakesh Chada, Zhaoheng Zheng, Pradeep Natarajan
We propose a self-supervised shared encoder model that achieves strong results on several visual, language and multimodal benchmarks while being data, memory and run-time efficient…