collaborators

7 papers

cs.CL2026

Vision-Language Models are Fragile Multilingual Associators

Ritabrata Chakraborty, Rajatsubhra Chakraborty, Shivakumara Palaiahnakote +2

Vision-language models must associate visual entities with textual attributes. Whether these associations or concept bindings remain stable when the language of the input changes i…

cs.CV2026

Exploring Mutual Cross-Modal Attention for Context-Aware Human Affordance Generation

Prasun Roy, Saumik Bhattacharya, Subhankar Ghosh +2

Human affordance learning investigates contextually relevant novel pose prediction such that the estimated pose represents a valid human action within the scene. While the task is…

cs.CV2026

TaleDiffusion: Multi-Character Story Generation with Dialogue Rendering

Ayan Banerjee, Josep Llados, Umapada Pal +1

Text-to-story visualization is challenging due to the need for consistent interaction among multiple characters across frames. Existing methods struggle with character consistency,…

cs.CV2025

CraftSVG: Multi-Object Text-to-SVG Synthesis via Layout Guided Diffusion

Ayan Banerjee, Nityanand Mathur, Josep Llados +2

Generating VectorArt from text prompts is a challenging vision task, requiring diverse yet realistic depictions of the seen as well as unseen entities. However, existing research h…

cs.CV2025

Privacy-enhancing Sclera Segmentation Benchmarking Competition: SSBC 2025

Matej Vitek, Darian Tomašević, Abhijit Das +32

This paper presents a summary of the 2025 Sclera Segmentation Benchmarking Competition (SSBC), which focused on the development of privacy-preserving sclera-segmentation models tra…

cs.CV2025

FastTextSpotter: A High-Efficiency Transformer for Multilingual Scene Text Spotting

Alloy Das, Sanket Biswas, Umapada Pal +2

The proliferation of scene text in both structured and unstructured environments presents significant challenges in optical character recognition (OCR), necessitating more efficien…