Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
-CLIP: Text-Conditioned Contrastive Learning for Multi-Granular Vision-Language Alignment
Fatimah Zohra, Chen Zhao, Hani Itani +1
CLIP achieves strong zero-shot image-text retrieval by aligning global vision and text representations, yet it falls behind on fine-grained tasks even when fine-tuned on long, deta…
cs.CV2024
SynthCLIP: Are We Ready for a Fully Synthetic CLIP Training?
Hasan Abed Al Kader Hammoud, Hani Itani, Fabio Pizzati +3
We present SynthCLIP, a CLIP model trained on entirely synthetic text-image pairs. Leveraging recent text-to-image (TTI) networks and large language models (LLM), we generate synth…
cs.CV2024
Compressed-Language Models for Understanding Compressed File Formats: a JPEG Exploration
Juan C. Pérez, Alejandro Pardo, Mattia Soldan +3
This study investigates whether Compressed-Language Models (CLMs), i.e. language models operating on raw byte streams from Compressed File Formats~(CFFs), can understand files comp…