MuG: A Multimodal Classification Benchmark on Game Data with Tabular, Textual, and Visual Fields
arXiv:2302.02978 · doi:10.18653/v1/2023.findings-emnlp.354
Abstract
Previous research has demonstrated the advantages of integrating data from multiple sources over traditional unimodal data, leading to the emergence of numerous novel multimodal applications. We propose a multimodal classification benchmark MuG with eight datasets that allows researchers to evaluate and improve their models. These datasets are collected from four various genres of games that cover tabular, textual, and visual modalities. We conduct multi-aspect data analysis to provide insights into the benchmark, including label balance ratios, percentages of missing features, distributions of data within each modality, and the correlations between labels and input modalities. We further present experimental results obtained by several state-of-the-art unimodal classifiers and multimodal classifiers, which demonstrate the challenging and multimodal-dependent properties of the benchmark. MuG is released at https://github.com/lujiaying/MUG-Bench with the data, tutorials, and implemented baselines.
References in corpus (11)
- Zero-Shot Text-to-Image Generation
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
- Visual Instruction Tuning
- DreamFusion: Text-to-3D using 2D Diffusion
- Kosmos-2: Grounding Multimodal Large Language Models to the World
- TabLLM: Few-shot Classification of Tabular Data with Large Language Models
- Emu: Generative Pretraining in Multimodality
- TabGNN: Multiplex Graph Neural Network for Tabular Data Prediction
- MedDiff: Generating Electronic Health Records using Accelerated Denoising Diffusion Model
- Reasoning about Actions over Visual and Linguistic Modalities: A Survey
- Evaluation and Enhancement of Semantic Grounding in Large Vision-Language Models