3 papers
cs.CV2026
SCI-CLIP: Segment-Centric Inference with Reference Memory for Training-Free Open-Vocabulary Segmentation
Mohamad Zamini, Diksha Shukla
Training-free open-vocabulary segmentation remains limited by a missing inference abstraction. Frozen vision-language features are produced at patch level, yet dense prediction req…
cs.CV2026
DouC: Dual-Branch CLIP for Training-Free Open-Vocabulary Segmentation
Mohamad Zamini, Diksha Shukla
Open-vocabulary semantic segmentation requires assigning pixel-level semantic labels while supporting an open and unrestricted set of categories. Training-free CLIP-based approache…
cs.CV2025
Delta-LLaVA: Base-then-Specialize Alignment for Token-Efficient Vision-Language Models
Mohamad Zamini, Diksha Shukla
Multimodal Large Language Models (MLLMs) combine visual and textual representations to enable rich reasoning capabilities. However, the high computational cost of processing dense…