11 papers · 1 filter
AGVBench: A Reliability-Oriented Benchmark of Data Augmentation for Vein Recognition
Haiyang Li, Yuming Fu, Qun Song +6
Vein recognition is a secure biometric technology often constrained by limited annotated data and imaging variations. While data augmentation mitigates this, strategies designed fo…
EarlyTom: Early Token Compression Completes Fast Video Understanding
Hesong Wang, Xin Jin, Lu Lu +4
Video large language models (Video-LLMs) have demonstrated strong capabilities in video understanding tasks. However, their practical deployment is still hindered by the inefficien…
LVOmniBench: Pioneering Long Audio-Video Understanding Evaluation for Omnimodal LLMs
Keda Tao, Yuhua Zheng, Jia Xu +13
Recent advancements in omnimodal large language models (OmniLLMs) have significantly improved the comprehension of audio and video inputs. However, current evaluations primarily fo…
MergeMix: A Unified Augmentation Paradigm for Visual and Multi-Modal Understanding
Xin Jin, Siyuan Li, Siyong Jian +2
Vision-language alignment in multi-modal large language models (MLLMs) relies on supervised fine-tuning (SFT) or reinforcement learning (RL). To align multi-modal large language mo…
MS-Mix: Sentiment-Guided Adaptive Augmentation for Multimodal Sentiment Analysis
Hongyu Zhu, Lin Chen, Xin Jin +1
Multimodal Sentiment Analysis (MSA) integrates complementary features from text, video, and audio for robust emotion understanding in human interactions. However, models suffer fro…
Relax DARTS: Relaxing the Constraints of Differentiable Architecture Search for Eye Movement Recognition
Hongyu Zhu, Xin Jin, Hongchao Liao +3
Eye movement biometrics is a secure and innovative identification method. Deep learning methods have shown good performance, but their network architecture relies on manual design…