5 papers
Is SAM3 ready for pathology segmentation?
Qiuyu Kong, Shakiba Sharifi, Yiming Wang +2
Is Segment Anything Model 3 (SAM3) capable in segmenting Any Pathology Images? Digital pathology segmentation spans tissue-level and nuclei-level scales, where traditional methods…
StructXLIP: Enhancing Vision-language Models with Multimodal Structural Cues
Zanxi Ruan, Songqun Gao, Qiuyu Kong +2
Edge-based representations are fundamental cues for visual understanding, a principle rooted in early vision research and still central today. We extend this principle to vision-la…
Multi-Level Conditioning by Pairing Localized Text and Sketch for Fashion Image Generation
Ziyue Liu, Davide Talon, Federico Girella +5
Sketches offer designers a concise yet expressive medium for early-stage fashion ideation by specifying structure, silhouette, and spatial relationships, while textual descriptions…
BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users
Wanyin Cheng, Zanxi Ruan
Visual Question Answering (VQA) holds great potential for assisting Blind and Low Vision (BLV) users, yet real-world usage remains challenging. Due to visual impairments, BLV users…
LOTS of Fashion! Multi-Conditioning for Image Generation via Sketch-Text Pairing
Federico Girella, Davide Talon, Ziyue Liu +3
Fashion design is a complex creative process that blends visual and textual expressions. Designers convey ideas through sketches, which define spatial structure and design elements…