activity
20242026
collaborators

15 papers

cs.CV2026

EVL-MCoT: Enhanced Vision-Language Multi-CoT for Harmful Meme Detection

Hao Yang, Jin Wang, Xuejie Zhang

MEMEs are widely used on the internet and often carry strong elements of sarcasm or irony. Understanding their hidden meanings typically requires a joint interpretation of text and…

cs.CV2026

Visual Saliency Steering Distillation for Multimodal Chain-of-Thought Reasoning

Hao Yang, Jin Wang, Xuejie Zhang

Multimodal chain-of-thought (CoT) reasoning integrates visual and textual cues through step-by-step inference. In small models with limited token budgets, modality-interaction fusi…

cs.CL2026

SAPO: Self-Adaptive Process Optimization Makes Small Reasoners Stronger

Kaiyuan Chen, Guangmin Zheng, Jin Wang +2

Existing self-evolution methods overlook the influence of fine-grained reasoning steps, which leads to the reasoner-verifier gap. The computational inefficiency of Monte Carlo (MC)…

cs.AI2026

LLMdoctor: Token-Level Flow-Guided Preference Optimization for Efficient Test-Time Alignment of Large Language Models

Tiesunlong Shen, Rui Mao, Jin Wang +4

Aligning Large Language Models (LLMs) with human preferences is critical, yet traditional fine-tuning methods are computationally expensive and inflexible. While test-time alignmen…

cs.CL2025

Sample-aware Adaptive Structured Pruning for Large Language Models

Jun Kong, Xinge Ma, Jin Wang +1

Large language models (LLMs) have achieved outstanding performance in natural language processing, but enormous model sizes and high computational costs limit their practical deplo…

cs.CV2025

Vision-aware Multimodal Prompt Tuning for Uploadable Multi-source Few-shot Domain Adaptation

Kuanghong Liu, Jin Wang, Kangjian He +2

Conventional multi-source domain few-shot adaptation (MFDA) faces the challenge of further reducing the load on edge-side devices in low-resource scenarios. Considering the native…