15 papers
EVL-MCoT: Enhanced Vision-Language Multi-CoT for Harmful Meme Detection
Hao Yang, Jin Wang, Xuejie Zhang
MEMEs are widely used on the internet and often carry strong elements of sarcasm or irony. Understanding their hidden meanings typically requires a joint interpretation of text and…
Visual Saliency Steering Distillation for Multimodal Chain-of-Thought Reasoning
Hao Yang, Jin Wang, Xuejie Zhang
Multimodal chain-of-thought (CoT) reasoning integrates visual and textual cues through step-by-step inference. In small models with limited token budgets, modality-interaction fusi…
SAPO: Self-Adaptive Process Optimization Makes Small Reasoners Stronger
Kaiyuan Chen, Guangmin Zheng, Jin Wang +2
Existing self-evolution methods overlook the influence of fine-grained reasoning steps, which leads to the reasoner-verifier gap. The computational inefficiency of Monte Carlo (MC)…
LLMdoctor: Token-Level Flow-Guided Preference Optimization for Efficient Test-Time Alignment of Large Language Models
Tiesunlong Shen, Rui Mao, Jin Wang +4
Aligning Large Language Models (LLMs) with human preferences is critical, yet traditional fine-tuning methods are computationally expensive and inflexible. While test-time alignmen…
Sample-aware Adaptive Structured Pruning for Large Language Models
Jun Kong, Xinge Ma, Jin Wang +1
Large language models (LLMs) have achieved outstanding performance in natural language processing, but enormous model sizes and high computational costs limit their practical deplo…
Vision-aware Multimodal Prompt Tuning for Uploadable Multi-source Few-shot Domain Adaptation
Kuanghong Liu, Jin Wang, Kangjian He +2
Conventional multi-source domain few-shot adaptation (MFDA) faces the challenge of further reducing the load on edge-side devices in low-resource scenarios. Considering the native…