activity
20242026
collaborators

7 papers

cs.CV2026

LVOmniBench: Pioneering Long Audio-Video Understanding Evaluation for Omnimodal LLMs

Keda Tao, Yuhua Zheng, Jia Xu +13

Recent advancements in omnimodal large language models (OmniLLMs) have significantly improved the comprehension of audio and video inputs. However, current evaluations primarily fo…

cs.CV2025

MergeMix: A Unified Augmentation Paradigm for Visual and Multi-Modal Understanding

Xin Jin, Siyuan Li, Siyong Jian +2

Vision-language alignment in multi-modal large language models (MLLMs) relies on supervised fine-tuning (SFT) or reinforcement learning (RL). To align multi-modal large language mo…

cs.CV2025

MS-Mix: Sentiment-Guided Adaptive Augmentation for Multimodal Sentiment Analysis

Hongyu Zhu, Lin Chen, Xin Jin +1

Multimodal Sentiment Analysis (MSA) integrates complementary features from text, video, and audio for robust emotion understanding in human interactions. However, models suffer fro…

cs.LG2025

Taming LLMs by Scaling Learning Rates with Gradient Grouping

Siyuan Li, Juanxi Tian, Zedong Wang +4

Training large language models (LLMs) poses challenges due to their massive scale and heterogeneous architectures. While adaptive optimizers like AdamW help address gradient variat…

cs.CV2024

Relax DARTS: Relaxing the Constraints of Differentiable Architecture Search for Eye Movement Recognition

Hongyu Zhu, Xin Jin, Hongchao Liao +3

Eye movement biometrics is a secure and innovative identification method. Deep learning methods have shown good performance, but their network architecture relies on manual design…

cs.CV2024

EM-DARTS: Hierarchical Differentiable Architecture Search for Eye Movement Recognition

Huafeng Qin, Hongyu Zhu, Xin Jin +3

Eye movement biometrics has received increasing attention thanks to its highly secure identification. Although deep learning (DL) models have shown success in eye movement recognit…