Large Language Models -- the Future of Fundamental Physics?
arXiv:2506.14757 · doi:10.21468/SciPostPhys.20.3.070
Abstract
For many fundamental physics applications, transformers, as the state of the art in learning complex correlations, benefit from pretraining on quasi-out-of-domain data. The obvious question is whether we can exploit Large Language Models, requiring proper out-of-domain transfer learning. We show how the Qwen2.5 LLM can be used to analyze and generate SKA data, specifically 3D maps of the cosmological large-scale structure for a large part of the observable Universe. We combine the LLM with connector networks and show, for cosmological parameter regression and lightcone generation, that this Lightcone LLM (L3M) with Qwen2.5 weights outperforms standard initialization and compares favorably with dedicated networks of matching size.
35 pages, 10 figures, 6 tables. v2: matched published version
References in corpus (88)
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- Deep Residual Learning for Image Recognition
- Training language models to follow instructions with human feedback
- LoRA: Low-Rank Adaptation of Large Language Models
- A Survey of Large Language Models
- Language Modeling with Gated Convolutional Networks
- The frontier of simulation-based inference
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- 21cmFAST: A Fast, Semi-Numerical Simulation of the High-Redshift 21-cm Signal
- A Survey on Multimodal Large Language Models
- Training Compute-Optimal Large Language Models
- Science with the Square Kilometer Array: Motivation, Key Science Projects, Standards and Assumptions
- A Comprehensive Overview of Large Language Models
- The CAMELS project: Cosmology and Astrophysics with MachinE Learning Simulations
- Hydrogen reionisation ends by : Lyman- optical depth measured by the XQR-30 sample
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- RoFormer: Enhanced Transformer with Rotary Position Embedding
- Large Language Models: A Survey
- Deep-learned Top Tagging with a Lorentz Layer
- Enhancing Gravitational-Wave Science with Machine Learning
- 21cmFAST v3: A Python-integrated C code forgenerating 3D realizations of the cosmic 21cm signal
- LIMA: Less Is More for Alignment
- An Efficient Lorentz Equivariant Graph Neural Network for Jet Tagging
- One Fits All:Power General Time Series Analysis by Pretrained LM
- Machine Learning and LHC Event Generation
- LLM4TS: Aligning Pre-Trained LLMs as Data-Efficient Time-Series Forecasters
- Qwen Technical Report
- Flow Matching for Generative Modeling
- Qwen2.5 Technical Report
- Root Mean Square Layer Normalization
- Large Language Models Are Zero-Shot Time Series Forecasters
- 21CMMC with a 3D light-cone: the impact of the co-evolution approximation on the astrophysics of reionisation and cosmic dawn
- New Constraints on Warm Dark Matter from the Lyman- Forest Power Spectrum
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Unveiling Dark Matter free-streaming at the smallest scales with high redshift Lyman-alpha forest
- Large Language Models for Forecasting and Anomaly Detection: A Systematic Literature Review
- Lorentz Group Equivariant Neural Network for Particle Physics
- Simulation-Based Inference of Reionization Parameters From 3D Tomographic 21 cm Lightcone Images
- AstroCLIP: A Cross-Modal Foundation Model for Galaxies
- Lattice gauge equivariant convolutional neural networks
- Qwen2 Technical Report
- Damping wings in the Lyman-α forest: a model-independent measurement of the neutral fraction at 5.4<z<6.1
- Symmetries, Safety, and Self-Supervision
- Particle Transformer for Jet Tagging
- Extending Context Window of Large Language Models via Positional Interpolation
- Exploring the likelihood of the 21-cm power spectrum with simulation-based inference
- Modern Machine Learning for LHC Physicists
- 21cmEMU: an emulator of 21cmFAST summary observables
- DropAttention: A Regularization Method for Fully-Connected Self-Attention Networks
- Investigating the Limitations of Transformers with Simple Arithmetic Tasks
- Learning the language of QCD jets with transformers
- Explainable Equivariant Neural Networks for Particle Physics: PELICAN
- GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
- How well do Large Language Models perform in Arithmetic tasks?
- OmniJet-: The first cross-task foundation model for particle physics
- Masked Particle Modeling on Sets: Towards Self-Supervised High Energy Physics Foundation Models
- Inferring Astrophysics and Dark Matter Properties from 21cm Tomography using Deep Learning
- Solving Key Challenges in Collider Physics with Foundation Models
- YaRN: Efficient Context Window Extension of Large Language Models
- Goat: Fine-tuned LLaMA Outperforms GPT-4 on Arithmetic Tasks
- Anomalies, Representations, and Self-Supervision
- Re-Simulation-based Self-Supervised Learning for Pre-Training Foundation Models
- A Lorentz-Equivariant Transformer for All of the LHC
- Machine Learning and Cosmology
- Lorentz-Equivariant Geometric Algebra Transformers for High-Energy Physics
- Model Reprogramming: Resource-Efficient Cross-Domain Machine Learning
- Optimal, fast, and robust inference of reionization-era cosmology with the 21cmPIE-INN
- Equivariant, Safe and Sensitive -- Graph Networks for New Physics
- Applications of Lattice Gauge Equivariant Neural Networks
- Teaching Arithmetic to Small Transformers
- Aspen Open Jets: Unlocking LHC Data for Foundation Models in Particle Physics
- SKATR: A Self-Supervised Summary Transformer for SKA
- Arithmetic with Language Models: from Memorization to Computation
- Optimal Equivariant Architectures from the Symmetries of Matrix-Element Likelihoods
- Equivariance and generalization in neural networks
- Is Tokenization Needed for Masked Particle Modelling?
- Bumblebee: Foundation Model for Particle Physics Discovery
- Finetuning Foundation Models for Joint Analysis Optimization
- AstroPT: Scaling Large Observation Models for Astronomy
- Toward a Spectral Foundation Model: An Attention-Based Approach with Domain-Inspired Fine-Tuning and Wavelength Parameterization
- Understanding LLM Embeddings for Regression
- Extrapolating Jet Radiation with Autoregressive Transformers
- Learning Broken Symmetries with Approximate Invariance
- Accelerating Resonance Searches via Signature-Oriented Pre-training
- HEP-JEPA: A foundation model for collider physics using joint embedding predictive architecture
- Maven: A Multimodal Foundation Model for Supernova Science
- Physics Event Classification Using Large Language Models
- The LoReLi database: 21 cm signal inference with 3D radiative hydrodynamics simulations