5 papers · 1 filter
VLRMBench: A Comprehensive and Challenging Benchmark for Vision-Language Reward Models
Jiacheng Ruan, Wenzhen Yuan, Xian Gao +6
Although large visual-language models (LVLMs) have demonstrated strong performance in multimodal tasks, errors may occasionally arise due to biases during the reasoning process. Re…
Understanding Robustness of Parameter-Efficient Tuning for Image Classification
Jiacheng Ruan, Xian Gao, Suncheng Xiang +3
Parameter-efficient tuning (PET) techniques calibrate the model's predictions on downstream tasks by freezing the pre-trained models and introducing a small number of learnable par…
MM-CamObj: A Comprehensive Multimodal Dataset for Camouflaged Object Scenarios
Jiacheng Ruan, Wenzhen Yuan, Zehao Lin +5
Large visual-language models (LVLMs) have achieved great success in multiple applications. However, they still encounter challenges in complex scenes, especially those involving ca…
Oceanship: A Large-Scale Dataset for Underwater Audio Target Recognition
Zeyu Li, Suncheng Xiang, Tong Yu +5
The recognition of underwater audio plays a significant role in identifying a vessel while it is in motion. Underwater target recognition tasks have a wide range of applications in…
iDAT: inverse Distillation Adapter-Tuning
Jiacheng Ruan, Jingsheng Gao, Mingye Xie +4
Adapter-Tuning (AT) method involves freezing a pre-trained model and introducing trainable adapter modules to acquire downstream knowledge, thereby calibrating the model for better…