2 papers
cs.CV2025
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor
Vatsal Agarwal, Matthew Gwilliam, Gefen Kohavi +3
Recent advances in multimodal large language models (MLLMs) have enabled image-based question-answering capabilities. However, a key limitation is the use of CLIP as the visual enc…
cs.CV2024
Revealing the Utilized Rank of Subspaces of Learning in Neural Networks
Isha Garg, Christian Koguchi, Eshan Verma +1
In this work, we study how well the learned weights of a neural network utilize the space available to them. This notion is related to capacity, but additionally incorporates the i…