53 citations · 140 across the 25 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
ImageNet-Think-250K: A Large-Scale Synthetic Dataset for Multimodal Reasoning for Vision Language Models
Krishna Teja Chitty-Venkata, Murali Emani
We develop ImageNet-Think, a multimodal reasoning dataset designed to aid the development of Vision Language Models (VLMs) with explicit reasoning capabilities. Our dataset is buil…
cs.CV2025
LangVision-LoRA-NAS: Neural Architecture Search for Variable LoRA Rank in Vision Language Models
Krishna Teja Chitty-Venkata, Murali Emani, Venkatram Vishwanath
Vision Language Models (VLMs) integrate visual and text modalities to enable multimodal understanding and generation. These models typically combine a Vision Transformer (ViT) as a…