4 papers · 1 filter
Seeing Beyond Redundancy: Task Complexity's Role in Vision Token Specialization in VLLMs
Darryl Hannan, John Cooper, Dylan White +1
Vision capabilities in vision large language models (VLLMs) have consistently lagged behind their linguistic capabilities. In particular, numerous benchmark studies have demonstrat…
Saddle-Free Guidance: Improved On-Manifold Sampling without Labels or Additional Training
Eric Yeats, Darryl Hannan, Wilson Fearn +3
Score-based generative models require guidance in order to generate plausible, on-manifold samples. The most popular guidance method, Classifier-Free Guidance (CFG), is only applic…
FMG-Det: Foundation Model Guided Robust Object Detection
Darryl Hannan, Timothy Doster, Henry Kvinge +2
Collecting high quality data for object detection tasks is challenging due to the inherent subjectivity in labeling the boundaries of an object. This makes it difficult to not only…
Foundation Models for Remote Sensing: An Analysis of MLLMs for Object Localization
Darryl Hannan, John Cooper, Dylan White +3
Multimodal large language models (MLLMs) have altered the landscape of computer vision, obtaining impressive results across a wide range of tasks, especially in zero-shot settings.…