7 citations · 7 across the 8 of their papers we have counts for
5 papers · 1 filter
Attribute-Conditioned Multimodal Slot Factorization for Controllable Fashion Retrieval
Najmeh Forouzandehmehr, Topojoy Biswas, Evren Korpeoglu +1
Fashion retrieval often requires satisfying multiple attributes at once, such as category, color, pattern, and demographic. Monolithic embeddings mix these signals into a single ve…
Segment and Matte Anything in a Unified Model
Zezhong Fan, Xiaohan Li, Topojoy Biswas +2
Segment Anything (SAM) has recently pushed the boundaries of segmentation by demonstrating zero-shot generalization and flexible prompting after training on over one billion masks.…
Spatial Reasoning in Foundation Models: Benchmarking Object-Centric Spatial Understanding
Vahid Mirjalili, Ramin Giahi, Sriram Kollipara +9
Spatial understanding is a critical capability for vision foundation models. While recent advances in large vision models or vision-language models (VLMs) have expanded recognition…
LayoutAgent: A Vision-Language Agent Guided Compositional Diffusion for Spatial Layout Planning
Zezhong Fan, Xiaohan Li, Luyi Ma +6
Designing realistic multi-object scenes requires not only generating images, but also planning spatial layouts that respect semantic relations and physical plausibility. On one han…
Prompt Optimizer of Text-to-Image Diffusion Models for Abstract Concept Understanding
Zezhong Fan, Xiaohan Li, Chenhao Fang +4
The rapid evolution of text-to-image diffusion models has opened the door of generative AI, enabling the translation of textual descriptions into visually compelling images with re…