222 citations · 328 across the 139 of their papers we have counts for
5 papers · 2 filters
Rice-VL: Evaluating Vision-Language Models for Cultural Understanding Across ASEAN Countries
Tushar Pranav, Eshan Pandey, Austria Lyka Diane Bala +3
Vision-Language Models (VLMs) excel in multimodal tasks but often exhibit Western-centric biases, limiting their effectiveness in culturally diverse regions like Southeast Asia (SE…
DeHate: A Stable Diffusion-based Multimodal Approach to Mitigate Hate Speech in Images
Dwip Dalal, Gautam Vashishtha, Anku Rani +11
The rise in harmful online content not only distorts public discourse but also poses significant challenges to maintaining a healthy digital environment. In response to this, we in…
Peccavi: Visual Paraphrase Attack Safe and Distortion Free Image Watermarking Technique for AI-Generated Images
Shreyas Dixit, Ashhar Aziz, Shashwat Bajpai +4
A report by the European Union Law Enforcement Agency predicts that by 2026, up to 90 percent of online content could be synthetically generated, raising concerns among policymaker…
DanceText: A Training-Free Layered Framework for Controllable Multilingual Text Transformation in Images
Zhenyu Yu, Mohd Yamani Idna Idris, Hua Wang +6
We present DanceText, a training-free framework for multilingual text editing in images, designed to support complex geometric transformations and achieve seamless foreground-backg…
From Fog to Failure: The Unintended Consequences of Dehazing on Object Detection in Clear Images
Ashutosh Kumar, Aman Chadha
This study explores the challenges of integrating human visual cue-based dehazing into object detection, given the selective nature of human perception. While human vision adapts d…