Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
Zero-Shot Vision Encoder Grafting via LLM Surrogates
Kaiyu Yue, Vasu Singla, Menglin Jia +6
Vision language models (VLMs) typically pair a modestly sized vision encoder with a large language model (LLM), e.g., Llama-70B, making the decoder the primary computational burden…
cs.CV2024
From Pixels to Prose: A Large Dataset of Dense Image Captions
Vasu Singla, Kaiyu Yue, Sukriti Paul +7
Training large vision-language models requires extensive, high-quality image-text pairs. Existing web-scraped datasets, however, are noisy and lack detailed image descriptions. To…