1 citations · 1 across the 5 of their papers we have counts for
2 papers
cs.CV2025
Reasoning under Vision: Understanding Visual-Spatial Cognition in Vision-Language Models for CAPTCHA
Python Song, Luke Tenyi Chang, Yun-Yun Tsai +2
CAPTCHA, originally designed to distinguish humans from robots, has evolved into a real-world benchmark for assessing the spatial reasoning capabilities of vision-language models.…
cs.CV2025
LAVID: An Agentic LVLM Framework for Diffusion-Generated Video Detection
Qingyuan Liu, Yun-Yun Tsai, Ruijian Zha +4
The impressive achievements of generative models in creating high-quality videos have raised concerns about digital integrity and privacy vulnerabilities. Recent works of AI-genera…