most citedEvoMoE: Expert Evolution in Mixture of Experts for Multimodal Large Language Models

1 citations · 1 across the 2 of their papers we have counts for

collaborators

5 papers

cs.CV2025

Are We Ready for RL in Text-to-3D Generation? A Progressive Investigation

Yiwen Tang, Zoey Guo, Kaixin Zhu +11

Reinforcement learning (RL), earlier proven to be effective in large language and multi-modal models, has been successfully extended to enhance 2D image generation recently. Howeve…

cs.CL20251 cited

EvoMoE: Expert Evolution in Mixture of Experts for Multimodal Large Language Models

Linglin Jing, Yuting Gao, Zhigang Wang +5

Recent advancements have shown that the Mixture of Experts (MoE) approach significantly enhances the capacity of large language models (LLMs) and improves performance on downstream…

cs.RO2025

Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Haoming Song, Delin Qu, Yuanqi Yao +9

Humans practice slow thinking before performing actual actions when handling complex tasks in the physical world. This thinking paradigm, recently, has achieved remarkable advancem…

cs.CV2025

AerialVG: A Challenging Benchmark for Aerial Visual Grounding by Exploring Positional Relations

Junli Liu, Qizhi Chen, Zhigang Wang +6

Visual grounding (VG) aims to localize target objects in an image based on natural language descriptions. In this paper, we propose AerialVG, a new task focusing on visual groundin…

cs.CV2025

Exploring the Potential of Encoder-free Architectures in 3D LMMs

Yiwen Tang, Zoey Guo, Zhuhao Wang +8

Encoder-free architectures have been preliminarily explored in the 2D Large Multimodal Models (LMMs), yet it remains an open question whether they can be effectively applied to 3D…