Publications (5)
Not all tokens contribute equally to diffusion learning
Guoqing Zhang, Lu Shi, Wanru Xu +4
With the rapid development of conditional diffusion models, significant progress has been made in text-to-video generation. However, we observe that these models often neglect sema…
MSD-Score: Multi-Scale Distributional Scoring for Reference-Free Image Caption Evaluation
Shichao Kan, Xuyang Zhang, Haojie Zhang +7
Evaluating image captions without references remains challenging because global embedding similarity often misses fine-grained mismatches such as hallucinated objects, missing attr…
Object Retrieval for Visual Question Answering with Outside Knowledge
Shichao Kan, Yuhai Deng, Jiale Fu +5
Retrieval-augmented generation (RAG) with large language models (LLMs) plays a crucial role in question answering, as LLMs possess limited knowledge and are not updated with contin…
Layer-adaptive Structured Pruning Guided by Latency
Siyuan Pan, Linna Zhang, Jie Zhang +3
Structured pruning can simplify network architecture and improve inference speed. Combined with the underlying hardware and inference engine in which the final model is deployed, b…
Patch Spatio-Temporal Relation Prediction for Video Anomaly Detection
Hao Shen, Lu Shi, Wanru Xu +3
Video Anomaly Detection (VAD), aiming to identify abnormalities within a specific context and timeframe, is crucial for intelligent Video Surveillance Systems. While recent deep le…