activity
20242026
most citedMegaPairs: Massive Data Synthesis For Universal Multimodal Retrieval

1 citations · 1 across the 5 of their papers we have counts for

collaborators

6 papers

cs.CV2026

EVA01: Unified Native 3D Understanding and Generation via Mixture-of-Transformers

Zongyuan Yang, Mingjing Yi, Wanli Ma +8

This paper addresses the challenge of integrating 3D meshes as a native modality within Multimodal Large Language Models (MLLMs). Diffusion-based large reconstruction models decoup…

cs.IR2025

MR-Bench: Going Beyond Matching to Reasoning in Multimodal Retrieval

Junjie Zhou, Ze Liu, Lei Xiong +10

Multimodal retrieval is becoming a crucial component of modern AI applications, yet its evaluation lags behind the demands of more realistic and challenging scenarios. Existing ben…

cs.CV20241 cited

MegaPairs: Massive Data Synthesis For Universal Multimodal Retrieval

Junjie Zhou, Zheng Liu, Ze Liu +6

Despite the rapidly growing demand for multimodal retrieval, progress in this field remains severely constrained by a lack of training data. In this paper, we introduce MegaPairs,…

cs.GR2024

DirectL: Efficient Radiance Fields Rendering for 3D Light Field Displays

Zongyuan Yang, Baolin Liu, Yingde Song +4

Autostereoscopic display, despite decades of development, has not achieved extensive application, primarily due to the daunting challenge of 3D content creation for non-specialists…

cs.IR2024

VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval

Junjie Zhou, Zheng Liu, Shitao Xiao +2

Multi-modal retrieval becomes increasingly popular in practice. However, the existing retrievers are mostly text-oriented, which lack the capability to process visual information.…

cs.CV2024

MLVU: Benchmarking Multi-task Long Video Understanding

Junjie Zhou, Yan Shu, Bo Zhao +9

The evaluation of Long Video Understanding (LVU) performance poses an important but challenging research problem. Despite previous efforts, the existing video understanding benchma…