activity
20212025
most citedSEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

52 citations · 161 across the 34 of their papers we have counts for

collaborators

29 papers

cs.CV2024

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation

Yuying Ge, Yizhuo Li, Yixiao Ge +1

In recent years, there has been a significant surge of interest in unifying image comprehension and generation within Large Language Models (LLMs). This growing interest has prompt…

cs.CV2024

ATP-LLaVA: Adaptive Token Pruning for Large Vision Language Models

Xubing Ye, Yukang Gan, Yixiao Ge +2

Large Vision Language Models (LVLMs) have achieved significant success across multi-modal tasks. However, the computational cost of processing long visual tokens can be prohibitive…

eess.SY2024

Geometric Data Fusion for Collaborative Attitude Estimation

Yixiao Ge, Behzad Zamani, Pieter van Goor +2

In this paper, we consider the collaborative attitude estimation problem for a multi-agent system. The agents are equipped with sensors that provide directional measurements and re…

cs.LG2024

GrootVL: Tree Topology is All You Need in State Space Model

Yicheng Xiao, Lin Song, Shaoli Huang +5

The state space models, employing recursively propagated features, demonstrate strong representation capabilities comparable to Transformer models and superior efficiency. However,…

cs.CV20243 cited

SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Bohao Li, Yuying Ge, Yi Chen +3

Comprehending text-rich visual content is paramount for the practical application of Multimodal Large Language Models (MLLMs), since text-rich scenarios are ubiquitous in the real…

cs.CV2024

ST-LLM: Large Language Models Are Effective Temporal Learners

Ruyang Liu, Chen Li, Haoran Tang +3

Large Language Models (LLMs) have showcased impressive capabilities in text comprehension and generation, prompting research efforts towards video LLMs to facilitate human-AI inter…