7 papers
Differentiable Efficient Operator Search
Xiaohuan Pei, Jiyuan Zhang, Yuanfan Guo +4
Efficient multimodal foundation models often rely on manually designed token-reduction operators, such as pruning, merging, pooling, and adaptive reweighting. Although these operat…
MoEBlaze: Breaking the Memory Wall for Efficient MoE Training on Modern GPUs
Jiyuan Zhang, Yining Liu, Siqi Yan +6
The pervasive "memory wall" bottleneck is significantly amplified in modern large-scale Mixture-of-Experts (MoE) architectures. MoE's inherent architectural sparsity leads to spars…
OmniSparse: Training-Aware Fine-Grained Sparse Attention for Long-Video MLLMs
Feng Chen, Yefei He, Shaoxuan He +9
Existing sparse attention methods primarily target inference-time acceleration by selecting critical tokens under predefined sparsity patterns. However, they often fail to bridge t…
Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding
Yun Li, Zhe Liu, Yajing Kong +6
Applying Multimodal Large Language Models (MLLMs) to video understanding presents significant challenges due to the need to model temporal relations across frames. Existing approac…
Automatic Item Generation for Personality Situational Judgment Tests with Large Language Models
Chang-Jin Li, Jiyuan Zhang, Yun Tang +1
Personality assessment through situational judgment tests (SJTs) offers unique advantages over traditional Likert-type self-report scales, yet their development remains labor-inten…
Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens
Feng Chen, Chenhui Gou, Jing Liu +6
As multimodal large language models (MLLMs) advance rapidly, rigorous evaluation has become essential, providing further guidance for their development. In this work, we focus on a…