11 papers
MAVISEG: Manifold Propagation and Visual Prototypes for Zero-Shot Open-Vocabulary Segmentation in Diffusion Transformers
Rajatsubhra Chakraborty, Xujun Che, Ritabrata Chakraborty +2
Text-to-image diffusion transformers learn about objects and scenes by learning to generate them, making them strong candidates for training-free zero-shot open-vocabulary semantic…
Vision-Language Models are Fragile Multilingual Associators
Ritabrata Chakraborty, Rajatsubhra Chakraborty, Shivakumara Palaiahnakote +2
Vision-language models must associate visual entities with textual attributes. Whether these associations or concept bindings remain stable when the language of the input changes i…
Andha-Dhun: A First Look at Audio Descriptions in Hindi
Ritabrata Chakraborty, Divy Kala, Nisheeth Bhooshan Gupta +3
Audio Descriptions (ADs) narrate visual content for Blind and Low Vision (BLV) audiences during gaps in audiovisual media. There is growing momentum around ADs in movies and TV sho…
A Lightweight Context-Driven Training-Free Network for Scene Text Segmentation and Recognition
Ritabrata Chakraborty, Shivakumara Palaiahnakote, Umapada Pal +1
Modern scene text recognition systems often depend on large end-to-end architectures that require extensive training and are prohibitively expensive for real-time scenarios. In suc…
Towards Robust Cross-Dataset Object Detection Generalization under Domain Specificity
Ritabrata Chakraborty, Hrishit Mitra, Shivakumara Palaiahnakote +1
Object detectors often perform well in-distribution, yet degrade sharply on a different benchmark. We study cross-dataset object detection (CD-OD) through a lens of setting specifi…
Do We Need Large VLMs for Spotting Soccer Actions?
Ritabrata Chakraborty, Rajatsubhra Chakraborty, Avijit Dasgupta +1
Traditional video-based tasks like soccer action spotting rely heavily on visual inputs, often requiring complex and computationally expensive models to process dense video data. W…