6 papers
EVL-MCoT: Enhanced Vision-Language Multi-CoT for Harmful Meme Detection
Hao Yang, Jin Wang, Xuejie Zhang
MEMEs are widely used on the internet and often carry strong elements of sarcasm or irony. Understanding their hidden meanings typically requires a joint interpretation of text and…
SAPO: Self-Adaptive Process Optimization Makes Small Reasoners Stronger
Kaiyuan Chen, Guangmin Zheng, Jin Wang +2
Existing self-evolution methods overlook the influence of fine-grained reasoning steps, which leads to the reasoner-verifier gap. The computational inefficiency of Monte Carlo (MC)…
Sample-aware Adaptive Structured Pruning for Large Language Models
Jun Kong, Xinge Ma, Jin Wang +1
Large language models (LLMs) have achieved outstanding performance in natural language processing, but enormous model sizes and high computational costs limit their practical deplo…
Vision-aware Multimodal Prompt Tuning for Uploadable Multi-source Few-shot Domain Adaptation
Kuanghong Liu, Jin Wang, Kangjian He +2
Conventional multi-source domain few-shot adaptation (MFDA) faces the challenge of further reducing the load on edge-side devices in low-resource scenarios. Considering the native…
Multi-Attribute Multi-Grained Adaptation of Pre-Trained Language Models for Text Understanding from Bayesian Perspective
You Zhang, Jin Wang, Liang-Chih Yu +2
Current neural networks often employ multi-domain-learning or attribute-injecting mechanisms to incorporate non-independent and identically distributed (non-IID) information for te…
Data-Free Black-Box Federated Learning via Zeroth-Order Gradient Estimation
Xinge Ma, Jin Wang, Xuejie Zhang
Federated learning (FL) enables decentralized clients to collaboratively train a global model under the orchestration of a central server without exposing their individual data. Ho…