6 papers
From Semantics, Scene to Instance-awareness: Distilling Foundation Model for Grounded Open-vocabulary Situation Recognition
Chen Cai, Tianyi Liu, Jianjun Gao +5
Recent Multimodal Large Language Models (MLLMs) exhibit strong zero-shot abilities but struggle with complex Grounded Situation Recognition (GSR) and are resource-intensive for edg…
PromptSR: Cascade Prompting for Lightweight Image Super-Resolution
Wenyang Liu, Chen Cai, Jianjun Gao +4
Although the lightweight Vision Transformer has significantly advanced image super-resolution (SR), it faces the inherent challenge of a limited receptive field due to the window-b…
SSH-Net: A Self-Supervised and Hybrid Network for Noisy Image Watermark Removal
Wenyang Liu, Jianjun Gao, Kim-Hui Yap
Visible watermark removal is challenging due to its inherent complexities and the noise carried within images. Existing methods primarily rely on supervised learning approaches tha…
Empowering Large Language Model for Continual Video Question Answering with Collaborative Prompting
Chen Cai, Zheng Wang, Jianjun Gao +4
In recent years, the rapid increase in online video content has underscored the limitations of static Video Question Answering (VideoQA) models trained on fixed datasets, as they s…
ByteNet: Rethinking Multimedia File Fragment Classification through Visual Perspectives
Wenyang Liu, Kejun Wu, Tianyi Liu +3
Multimedia file fragment classification (MFFC) aims to identify file fragment types, e.g., image/video, audio, and text without system metadata. It is of vital importance in multim…
CL-HOI: Cross-Level Human-Object Interaction Distillation from Vision Large Language Models
Jianjun Gao, Chen Cai, Ruoyu Wang +4
Human-object interaction (HOI) detection has seen advancements with Vision Language Models (VLMs), but these methods often depend on extensive manual annotations. Vision Large Lang…