7 papers
Kalman Filter Enhanced GRPO for Reinforcement Learning-Based Language Model Reasoning
Hu Wang, Congbo Ma, Ian Reid +1
The advantage function is a central concept in RL that helps reduce variance in policy gradient estimates. For language modeling, Group Relative Policy Optimization (GRPO) was prop…
TransPrune: Token Transition Pruning for Efficient Large Vision-Language Model
Ao Li, Yuxiang Duan, Jinghui Zhang +5
Large Vision-Language Models (LVLMs) have advanced multimodal learning but face high computational costs due to the large number of visual tokens, motivating token pruning to impro…
Meta-Learned Modality-Weighted Knowledge Distillation for Robust Multi-Modal Learning with Missing Data
Hu Wang, Salma Hassan, Yuyuan Liu +12
In multi-modal learning, some modalities are more influential than others, and their absence can have a significant impact on classification/segmentation accuracy. Addressing this…
Rethinking Weight-Averaged Model-merging
Hu Wang, Congbo Ma, Ibrahim Almakky +3
Model merging, particularly through weight averaging, has shown surprising effectiveness in saving computations and improving model performance without any additional training. How…
In-Model Merging for Enhancing the Robustness of Medical Imaging Classification Models
Hu Wang, Ibrahim Almakky, Congbo Ma +2
Model merging is an effective strategy to merge multiple models for enhancing model performances, and more efficient than ensemble learning as it will not introduce extra computati…
Learnable Cross-modal Knowledge Distillation for Multi-modal Learning with Missing Modality
Hu Wang, Congbo Ma, Jianpeng Zhang +4
The problem of missing modalities is both critical and non-trivial to be handled in multi-modal models. It is common for multi-modal tasks that certain modalities contribute more c…