10 papers
OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers
Siyuan Li, Jiabao Pan, Yumou Liu +9
Optimizer selection for large-scale model training has become a system-level design decision constrained jointly by compute, memory, tuning budget, and task diversity, yet the land…
AGVBench: A Reliability-Oriented Benchmark of Data Augmentation for Vein Recognition
Haiyang Li, Yuming Fu, Qun Song +6
Vein recognition is a secure biometric technology often constrained by limited annotated data and imaging variations. While data augmentation mitigates this, strategies designed fo…
EarlyTom: Early Token Compression Completes Fast Video Understanding
Hesong Wang, Xin Jin, Lu Lu +4
Video large language models (Video-LLMs) have demonstrated strong capabilities in video understanding tasks. However, their practical deployment is still hindered by the inefficien…
MS-Mix: Sentiment-Guided Adaptive Augmentation for Multimodal Sentiment Analysis
Hongyu Zhu, Lin Chen, Xin Jin +1
Multimodal Sentiment Analysis (MSA) integrates complementary features from text, video, and audio for robust emotion understanding in human interactions. However, models suffer fro…
LVOmniBench: Pioneering Long Audio-Video Understanding Evaluation for Omnimodal LLMs
Keda Tao, Yuhua Zheng, Jia Xu +13
Recent advancements in omnimodal large language models (OmniLLMs) have significantly improved the comprehension of audio and video inputs. However, current evaluations primarily fo…
MergeMix: A Unified Augmentation Paradigm for Visual and Multi-Modal Understanding
Xin Jin, Siyuan Li, Siyong Jian +2
Vision-language alignment in multi-modal large language models (MLLMs) relies on supervised fine-tuning (SFT) or reinforcement learning (RL). To align multi-modal large language mo…