activity
20242026
collaborators

10 papers

cs.CV2026

Is SAM3 ready for pathology segmentation?

Qiuyu Kong, Shakiba Sharifi, Yiming Wang +2

Is Segment Anything Model 3 (SAM3) capable in segmenting Any Pathology Images? Digital pathology segmentation spans tissue-level and nuclei-level scales, where traditional methods…

cs.CV2026

Anomaly-Aware Vision-Language Adapters for Zero-Shot Anomaly Detection

Muhammad Aqeel, Maham Nazir, Uzair Khan +2

Zero-shot anomaly detection aims to identify defects in unseen categories without target-specific training. Existing methods usually apply the same feature transformation to all sa…

cs.CV2026

Benchmarking Interaction, Beyond Policy: a Reproducible Benchmark for Collaborative Instance Object Navigation

Edoardo Zorzi, Francesco Taioli, Yiming Wang +4

We propose Question-Asking Navigation (QAsk-Nav), the first reproducible benchmark for Collaborative Instance Object Navigation (CoIN) that enables an explicit, separate assessment…

cs.CV2026

StructXLIP: Enhancing Vision-language Models with Multimodal Structural Cues

Zanxi Ruan, Songqun Gao, Qiuyu Kong +2

Edge-based representations are fundamental cues for visual understanding, a principle rooted in early vision research and still central today. We extend this principle to vision-la…

cs.CV2026

Multi-Level Conditioning by Pairing Localized Text and Sketch for Fashion Image Generation

Ziyue Liu, Davide Talon, Federico Girella +5

Sketches offer designers a concise yet expressive medium for early-stage fashion ideation by specifying structure, silhouette, and spatial relationships, while textual descriptions…

cs.CV2026

VISOR: VIsual Spatial Object Reasoning for Language-driven Object Navigation

Francesco Taioli, Shiping Yang, Sonia Raychaudhuri +3

Language-driven object navigation requires agents to interpret natural language descriptions of target objects, which combine intrinsic and extrinsic attributes for instance recogn…