4 papers
LUQ: Layerwise Ultra-Low Bit Quantization for Multimodal Large Language Models
Shubhang Bhatnagar, Andy Xu, Kar-Han Tan +1
Large Language Models (LLMs) with multimodal capabilities have revolutionized vision-language tasks, but their deployment often requires huge memory and computational resources. Po…
Beyond Scalar Distances: Semantic Attribute Gradients from Frozen MLLMs for Visual Embeddings
Shubhang Bhatnagar, Dheeraj Baiju, Narendra Ahuja
Vision encoders for retrieval are typically trained with class-label supervision: each training pair reduces to a scalar that uniformly pushes the embedding apart or pulls it toget…
Finding Distributed Object-Centric Properties in Self-Supervised Transformers
Samyak Rawlekar, Amitabh Swain, Yujun Cai +3
Self-supervised Vision Transformers (ViTs) like DINO show an emergent ability to discover objects, typically observed in [CLS] token attention maps of the final layer. However, the…
Potential Field Based Deep Metric Learning
Shubhang Bhatnagar, Narendra Ahuja
Deep metric learning (DML) involves training a network to learn a semantically meaningful representation space. Many current approaches mine n-tuples of examples and model interact…