1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.CL2025
Aleph-Alpha-GermanWeb: Improving German-language LLM pre-training with model-based data curation and synthetic data generation
Thomas F Burns, Letitia Parcalabescu, Stephan Wäldchen +5
Scaling data quantity is essential for large language models (LLMs), yet recent findings show that data quality can significantly boost performance and training efficiency. We intr…
econ.EM2025★ 1 cited
RUM-NN: A Neural Network Model Compatible with Random Utility Maximisation for Discrete Choice Setups
Niousha Bagheri, Milad Ghasri, Michael Barlow
This paper introduces a framework for capturing stochasticity of choice probabilities in neural networks, derived from and fully consistent with the Random Utility Maximization (RU…