activity
20242026
collaborators

11 papers

cs.CV2026

MAVISEG: Manifold Propagation and Visual Prototypes for Zero-Shot Open-Vocabulary Segmentation in Diffusion Transformers

Rajatsubhra Chakraborty, Xujun Che, Ritabrata Chakraborty +2

Text-to-image diffusion transformers learn about objects and scenes by learning to generate them, making them strong candidates for training-free zero-shot open-vocabulary semantic…

cs.CL2026

Vision-Language Models are Fragile Multilingual Associators

Ritabrata Chakraborty, Rajatsubhra Chakraborty, Shivakumara Palaiahnakote +2

Vision-language models must associate visual entities with textual attributes. Whether these associations or concept bindings remain stable when the language of the input changes i…

cs.CV2026

Andha-Dhun: A First Look at Audio Descriptions in Hindi

Ritabrata Chakraborty, Divy Kala, Nisheeth Bhooshan Gupta +3

Audio Descriptions (ADs) narrate visual content for Blind and Low Vision (BLV) audiences during gaps in audiovisual media. There is growing momentum around ADs in movies and TV sho…

cs.CV2026

A Lightweight Context-Driven Training-Free Network for Scene Text Segmentation and Recognition

Ritabrata Chakraborty, Shivakumara Palaiahnakote, Umapada Pal +1

Modern scene text recognition systems often depend on large end-to-end architectures that require extensive training and are prohibitively expensive for real-time scenarios. In suc…

cs.CV2026

Towards Robust Cross-Dataset Object Detection Generalization under Domain Specificity

Ritabrata Chakraborty, Hrishit Mitra, Shivakumara Palaiahnakote +1

Object detectors often perform well in-distribution, yet degrade sharply on a different benchmark. We study cross-dataset object detection (CD-OD) through a lens of setting specifi…

cs.CV2025

Do We Need Large VLMs for Spotting Soccer Actions?

Ritabrata Chakraborty, Rajatsubhra Chakraborty, Avijit Dasgupta +1

Traditional video-based tasks like soccer action spotting rely heavily on visual inputs, often requiring complex and computationally expensive models to process dense video data. W…