activity
20232026
most citedMoE-LLaVA: Mixture of Experts for Large Vision-Language Models

34 citations · 70 across the 10 of their papers we have counts for

collaborators
Showing cs.CVShow all

8 papers · 1 filter

cs.CV2026

SoPE: Spherical Coordinate-Based Positional Embedding for Enhancing Spatial Perception of 3D LVLMs

Guanting Ye, Qiyan Zhao, Wenhao Yu +7

3D Large Vision-Language Models (3D LVLMs) built upon Large Language Models (LLMs) have achieved remarkable progress across various multimodal tasks. However, their inherited posit…

cs.CV2025

LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs

Xiaodong Wang, Jinfa Huang, Li Yuan +1

Most Video Large Language Models (Video-LLMs) adopt preference alignment techniques, e.g., DPO~\citep{rafailov2024dpo}, to optimize the reward margin between a winning response ($y…

cs.CV2024★ 1 cited

RAP: Efficient Text-Video Retrieval with Sparse-and-Correlated Adapter

Meng Cao, Haoran Tang, Jinfa Huang +7

Text-Video Retrieval (TVR) aims to align relevant video content with natural language queries. To date, most state-of-the-art TVR methods learn image-to-video transfer learning bas…

cs.CV2024★ 34 cited

MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

Bin Lin, Zhenyu Tang, Yang Ye +7

Recent advances demonstrate that scaling Large Vision-Language Models (LVLMs) effectively improves downstream task performances. However, existing scaling methods enable all model…

cs.CV2023★ 5 cited

Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Munan Ning, Bin Zhu, Yujia Xie +5

Video-based large language models (Video-LLMs) have been recently introduced, targeting both fundamental improvements in perception and comprehension, and a diverse range of user i…

cs.CV2023★ 2 cited

Act As You Wish: Fine-Grained Control of Motion Diffusion Model with Hierarchical Semantic Graphs

Peng Jin, Yang Wu, Yanbo Fan +3

Most text-driven human motion generation methods employ sequential modeling approaches, e.g., transformer, to extract sentence-level text representations automatically and implicit…