240 citations · 830 across the 189 of their papers we have counts for
151 papers · 1 filter
Invisible in Space, Visible in Time: Motion Vision CAPTCHA against GUI Agents
Zeyu Zhang, Dingyi Rong, Zijian Chen +3
Most existing visual CAPTCHAs remain spatially solvable: the required information is exposed by static appearance, local structure, and interface state. This assumption is weakened…
LLaVA-Assessor: Building the Foundation LMM For Visual Quality Assessment
Ziheng Jia, Zicheng Zhang, Jiaying Qian +2
Aligning with the human visual system~(HVS) in perceiving and evaluating the quality of visual signals is a central objective of machine-vision-based visual quality assessment syst…
Emo-Bench: A Scalable Benchmark for Multimodal Evoked and Expressed Emotion Understanding via Bayesian Pairwise Alignment
Lancheng Gao, Ziheng Jia, Shengyan Li +4
Understanding both expressed and evoked emotions is critical for multimodal large language models (MLLMs) to achieve comprehensive affect-aware interactions. However, existing benc…
Visual Distortion Detection in UGC Images Using Large Multimodal Models
Ziheng Jia, Yingji Liang, Jiaying Qian +1
The localized depiction of perceptual quality has long been a crucial, yet underexplored, challenge in image quality assessment (IQA). Existing approaches based on large multimodal…
MIEScore: Human-Aligned Evaluation for Multi-Source Image Editing
Zitong Xu, Huiyu Duan, Xinyun Zhang +7
Recent advances in unified multimodal models have significantly improved text-guided image editing abilities. In particular, models such as Nano-Banana-Pro and GPT-Image-2 demonstr…
Multi-Dimensional Quality Assessment for AI-Generated Human-Centric Videos: Dataset and Model
Sijing Wu, Yunhao Li, Huiyu Duan +4
AI-generated human-centric videos play a crucial role in a wide range of modern applications. However, they often suffer from quality issues and semantic mismatches, underscoring t…