2 papers
cs.CV2025
Mimicking Human Visual Development for Learning Robust Image Representations
Ankita Raj, Kaashika Prajaapat, Tapan Kumar Gandhi +1
The human visual system is remarkably adept at adapting to changes in the input distribution; a capability modern convolutional neural networks (CNNs) still struggle to match. Draw…
cs.CV2025
CrossMed: A Multimodal Cross-Task Benchmark for Compositional Generalization in Medical Imaging
Pooja Singh, Siddhant Ujjain, Tapan Kumar Gandhi +1
Recent advances in multimodal large language models have enabled unified processing of visual and textual inputs, offering promising applications in general-purpose medical AI. How…