7 papers
Advancing Vision Transformer with Enhanced Spatial Priors
Qihang Fan, Huaibo Huang, Mingrui Chen +2
In recent years, the Vision Transformer (ViT) has garnered significant attention within the computer vision community. However, the core component of ViT, Self-Attention, lacks exp…
Think 360°: Evaluating the Width-centric Reasoning Capability of MLLMs Beyond Depth
Mingrui Chen, Hexiong Yang, Haogeng Liu +2
In this paper, we present a holistic multimodal benchmark that evaluates the reasoning capabilities of MLLMs with an explicit focus on reasoning width, a complementary dimension to…
Unlocking the Potential of Difficulty Prior in RL-based Multimodal Reasoning
Mingrui Chen, Haogeng Liu, Hao Liang +3
In this work, we investigate how explicitly modeling problem's difficulty prior information shapes the effectiveness of reinforcement learning based fine-tuning for multimodal reas…
Binary and Ternary Quantization Can Enhance Feature Discrimination
Weizhi Lu, Mingrui Chen, Weiyu Li
Quantization is widely applied in machine learning to reduce computational and storage costs for both data and models. Considering that classification tasks are fundamental to the…
Semantic Equitable Clustering: A Simple and Effective Strategy for Clustering Vision Tokens
Qihang Fan, Huaibo Huang, Mingrui Chen +1
The Vision Transformer (ViT) has gained prominence for its superior relational modeling prowess. However, its global attention mechanism's quadratic complexity poses substantial co…
HAD: Hybrid Architecture Distillation Outperforms Teacher in Genomic Sequence Modeling
Hexiong Yang, Mingrui Chen, Huaibo Huang +4
Inspired by the great success of Masked Language Modeling (MLM) in the natural language domain, the paradigm of self-supervised pre-training and fine-tuning has also achieved remar…