6 papers
V-FIND: Revealing the Intrinsic Forgery Knowledge Encoded in Video Forgery Detectors
Shichao Kan, Chengpeng Hong, Jingtong Dou +8
As generated videos become increasingly realistic, reliable video forgery detection is increasingly important. Existing studies typically optimize and use video forgery detectors a…
MSD-Score: Multi-Scale Distributional Scoring for Reference-Free Image Caption Evaluation
Shichao Kan, Xuyang Zhang, Haojie Zhang +7
Evaluating image captions without references remains challenging because global embedding similarity often misses fine-grained mismatches such as hallucinated objects, missing attr…
Dynamic Residual Encoding with Slide-Level Contrastive Learning for End-to-End Whole Slide Image Representation
Jing Jin, Xu Liu, Te Gao +6
Whole Slide Image (WSI) representation is critical for cancer subtyping, cancer recognition and mutation prediction.Training an end-to-end WSI representation model poses significan…
Contrastive Regularization over LoRA for Multimodal Biomedical Image Incremental Learning
Haojie Zhang, Yixiong Liang, Hulin Kuang +5
Multimodal Biomedical Image Incremental Learning (MBIIL) is essential for handling diverse tasks and modalities in the biomedical domain, as training separate models for each modal…
Object Retrieval for Visual Question Answering with Outside Knowledge
Shichao Kan, Yuhai Deng, Jiale Fu +5
Retrieval-augmented generation (RAG) with large language models (LLMs) plays a crucial role in question answering, as LLMs possess limited knowledge and are not updated with contin…
HRDecoder: High-Resolution Decoder Network for Fundus Image Lesion Segmentation
Ziyuan Ding, Yixiong Liang, Shichao Kan +1
High resolution is crucial for precise segmentation in fundus images, yet handling high-resolution inputs incurs considerable GPU memory costs, with diminishing performance gains a…