3 papers
cs.CV2026
DiVE-k: Differential Visual Reasoning for Fine-grained Image Recognition
Raja Kumar, Arka Sadhu, Ram Nevatia
Large Vision Language Models (LVLMs) possess extensive text knowledge but struggles to utilize this knowledge for fine-grained image recognition, often failing to differentiate bet…
cs.CV2025
Harnessing Shared Relations via Multimodal Mixup Contrastive Learning for Multimodal Classification
Raja Kumar, Raghav Singhal, Pranamya Kulkarni +2
Deep multimodal learning has shown remarkable success by leveraging contrastive learning to capture explicit one-to-one relations across modalities. However, real-world data often…
cs.CV2024
Few-shot Novel View Synthesis using Depth Aware 3D Gaussian Splatting
Raja Kumar, Vanshika Vats
3D Gaussian splatting has surpassed neural radiance field methods in novel view synthesis by achieving lower computational costs and real-time high-quality rendering. Although it p…