2 papers
cs.LG2025
Quantifying Modality Contributions via Disentangling Multimodal Representations
Padegal Amit, Omkar Mahesh Kashyap, Namitha Rayasam +2
Quantifying modality contributions in multimodal models remains a challenge, as existing approaches conflate the notion of contribution itself. Prior work relies on accuracy-based…
cs.CV2024
Enhancing Vision Models for Text-Heavy Content Understanding and Interaction
Adithya TG, Adithya SK, Abhinav R Bharadwaj +2
Interacting and understanding with text heavy visual content with multiple images is a major challenge for traditional vision models. This paper is on enhancing vision models' capa…