activity
20242026
most citedInfiMM-WebMath-40B: Advancing Multimodal Pre-Training for Enhanced Mathematical Reasoning

1 citations · 1 across the 9 of their papers we have counts for

collaborators
Showing cs.CVShow all

11 papers · 1 filter

cs.CV2026

Advancing Vision Transformer with Enhanced Spatial Priors

Qihang Fan, Huaibo Huang, Mingrui Chen +2

In recent years, the Vision Transformer (ViT) has garnered significant attention within the computer vision community. However, the core component of ViT, Self-Attention, lacks exp…

cs.CV2026

MVPBench: A Multi-Video Perception Evaluation Benchmark for Multi-Modal Video Understanding

Purui Bai, Tao Wu, Jiayang Sun +3

The rapid progress of Large Language Models (LLMs) has spurred growing interest in Multi-modal LLMs (MLLMs) and motivated the development of benchmarks to evaluate their perceptual…

cs.CV2026

Think 360°: Evaluating the Width-centric Reasoning Capability of MLLMs Beyond Depth

Mingrui Chen, Hexiong Yang, Haogeng Liu +2

In this paper, we present a holistic multimodal benchmark that evaluates the reasoning capabilities of MLLMs with an explicit focus on reasoning width, a complementary dimension to…

cs.CV2026

Random Wins All: Rethinking Grouping Strategies for Vision Tokens

Qihang Fan, Yuang Ai, Huaibo Huang +1

Since Transformers are introduced into vision architectures, their quadratic complexity has always been a significant issue that many research efforts aim to address. A representat…

cs.CV2025

Expand and Prune: Maximizing Trajectory Diversity for Effective GRPO in Generative Models

Shiran Ge, Chenyi Huang, Yuang Ai +3

Group Relative Policy Optimization (GRPO) is a powerful technique for aligning generative models, but its effectiveness is bottlenecked by the conflict between large group sizes an…

cs.CV2025

Rectifying Magnitude Neglect in Linear Attention

Qihang Fan, Huaibo Huang, Yuang Ai +1

As the core operator of Transformers, Softmax Attention exhibits excellent global modeling capabilities. However, its quadratic complexity limits its applicability to vision tasks.…