7 papers
Vision-Language Models are Fragile Multilingual Associators
Ritabrata Chakraborty, Rajatsubhra Chakraborty, Shivakumara Palaiahnakote +2
Vision-language models must associate visual entities with textual attributes. Whether these associations or concept bindings remain stable when the language of the input changes i…
Exploring Mutual Cross-Modal Attention for Context-Aware Human Affordance Generation
Prasun Roy, Saumik Bhattacharya, Subhankar Ghosh +2
Human affordance learning investigates contextually relevant novel pose prediction such that the estimated pose represents a valid human action within the scene. While the task is…
TaleDiffusion: Multi-Character Story Generation with Dialogue Rendering
Ayan Banerjee, Josep Llados, Umapada Pal +1
Text-to-story visualization is challenging due to the need for consistent interaction among multiple characters across frames. Existing methods struggle with character consistency,…
CraftSVG: Multi-Object Text-to-SVG Synthesis via Layout Guided Diffusion
Ayan Banerjee, Nityanand Mathur, Josep Llados +2
Generating VectorArt from text prompts is a challenging vision task, requiring diverse yet realistic depictions of the seen as well as unseen entities. However, existing research h…
Privacy-enhancing Sclera Segmentation Benchmarking Competition: SSBC 2025
Matej Vitek, Darian TomaÅ¡eviÄ, Abhijit Das +32
This paper presents a summary of the 2025 Sclera Segmentation Benchmarking Competition (SSBC), which focused on the development of privacy-preserving sclera-segmentation models tra…
FastTextSpotter: A High-Efficiency Transformer for Multilingual Scene Text Spotting
Alloy Das, Sanket Biswas, Umapada Pal +2
The proliferation of scene text in both structured and unstructured environments presents significant challenges in optical character recognition (OCR), necessitating more efficien…