7 citations · 51 across the 110 of their papers we have counts for
16 papers · 1 filter
GuardPaint:SpeculativeSafetyDecodingforText-to-ImageGeneration
Shreyash Dhoot, Paras Dhiman, Arsh Abbas Naqvi +4
Text-to-image (T2I) diffusion models offer powerful visual generation, but their controllability creates a critical safety challenge: adversarial prompts can steer the denoising tr…
VISTA: Video Interaction Spatio-Temporal Analysis Benchmark
Alejandro Aparcedo, Akash Kumar, Aaryan Garg +5
Existing benchmarks for Vision-Language Models (VLMs) primarily evaluate spatio-temporal understanding on simple single-action videos, closed attribute sets and restricted entity t…
SleepWalk: A Three-Tier Benchmark for Stress-Testing Instruction-Guided Vision-Language Navigation
Niyati Rawal, Sushant Ravva, Shah Alam Abir +5
Vision-Language Models (VLMs) have advanced rapidly in multimodal perception and language understanding, yet it remains unclear whether they can reliably ground language into spati…
Findings of the Counter Turing Test: AI-Generated Image Detection
Rajarshi Roy, Nasrin Imanpour, Ashhar Aziz +16
The rapid advancements in generative AI technologies, such as Stable Diffusion, DALL-E, and Midjourney, have significantly transformed the creation of synthetic visual content. Whi…
A Comprehensive Dataset for Human vs. AI Generated Image Detection
Rajarshi Roy, Ashhar Aziz, Shashwat Bajpai +17
Multimodal generative AI systems like Stable Diffusion, DALL-E, and MidJourney have fundamentally changed how synthetic images are created. These tools drive innovation but also en…
Rice-VL: Evaluating Vision-Language Models for Cultural Understanding Across ASEAN Countries
Tushar Pranav, Eshan Pandey, Austria Lyka Diane Bala +3
Vision-Language Models (VLMs) excel in multimodal tasks but often exhibit Western-centric biases, limiting their effectiveness in culturally diverse regions like Southeast Asia (SE…