The 2024 MRSI Data Processing and Quantification Challenge Synthetic Dataset
arXiv:2609.02637
Abstract
Synthetic data is central to magnetic resonance spectroscopy method development because they provide ground truths for software validation, reproducible benchmarking, and machine- and deep-learning training. In MRSI, synthetic data must capture spatially varying anatomy, field inhomogeneity, nuisance signals, and measurement effects. The 2024 MRSI Data Processing and Quantification Challenge Synthetic Dataset is a simulated 3T brain FID-MRSI resource with ground-truth metabolite maps developed as a controlled testbed for MRSI processing and quantification methods. Subject-specific simulations used anatomical images and field maps from Human Connectome Project subjects. Tissue masks, quantum-mechanically simulated metabolite basis functions, in vivo-derived macromolecular components, Bloch-simulated post-WET residual water, registered in vivo lipid signals, spectral baseline, -dependent frequency shifts, Voigt lineshape variations, and complex Gaussian noise were combined in a forward model. Water and lipid signals were synthesized on a high-resolution grid and Fourier-truncated to the final echo-planar spectroscopic imaging grid to model the finite spatial point spread function. This resource has 24 training and 8 testing datasets containing contaminated FID-MRSI data, anatomical images, maps, metadata, and ground-truth metabolite and component signals. Forward-model parameters are documented, including tissue-specific metabolite concentrations and relaxation times, macromolecular amplitude ratios, the water model, noise, and field-inhomogeneity ranges. This synthetic benchmark supports development and comparison of MRSI processing, nuisance-signal removal, and metabolite-quantification methods, while enabling method evaluation. It illustrates how characterized synthetic data can support rigorous evaluation when ground truth is difficult or impossible to obtain.
21 pages, 9 figures, 9 tables