6 citations · 13 across the 14 of their papers we have counts for
Showing 2025 · cs.CVShow all
2 papers · 2 filters
cs.CV2025
MergeMix: A Unified Augmentation Paradigm for Visual and Multi-Modal Understanding
Xin Jin, Siyuan Li, Siyong Jian +2
Vision-language alignment in multi-modal large language models (MLLMs) relies on supervised fine-tuning (SFT) or reinforcement learning (RL). To align multi-modal large language mo…
cs.CV2025
MS-Mix: Sentiment-Guided Adaptive Augmentation for Multimodal Sentiment Analysis
Hongyu Zhu, Lin Chen, Xin Jin +1
Multimodal Sentiment Analysis (MSA) integrates complementary features from text, video, and audio for robust emotion understanding in human interactions. However, models suffer fro…