1 citations · 1 across the 2 of their papers we have counts for
3 papers
cs.CV2026
MSD-Score: Multi-Scale Distributional Scoring for Reference-Free Image Caption Evaluation
Shichao Kan, Xuyang Zhang, Haojie Zhang +7
Evaluating image captions without references remains challenging because global embedding similarity often misses fine-grained mismatches such as hallucinated objects, missing attr…
cs.CV2024★ 1 cited
Patch Spatio-Temporal Relation Prediction for Video Anomaly Detection
Hao Shen, Lu Shi, Wanru Xu +3
Video Anomaly Detection (VAD), aiming to identify abnormalities within a specific context and timeframe, is crucial for intelligent Video Surveillance Systems. While recent deep le…
cs.CV2024
Object Retrieval for Visual Question Answering with Outside Knowledge
Shichao Kan, Yuhai Deng, Jiale Fu +5
Retrieval-augmented generation (RAG) with large language models (LLMs) plays a crucial role in question answering, as LLMs possess limited knowledge and are not updated with contin…