3 papers
cs.CL2026
Repetition over Diversity: High-Signal Data Filtering for Sample-Efficient German Language Modeling
Ansar Aynetdinov, Patrick Haller, Alan Akbik
Recent research has shown that filtering massive English web corpora into high-quality subsets significantly improves training efficiency. However, for high-resource non-English la…
cs.CL2025
Pre-Training Curriculum for Multi-Token Prediction in Language Models
Ansar Aynetdinov, Alan Akbik
Multi-token prediction (MTP) is a recently proposed pre-training objective for language models. Rather than predicting only the next token (NTP), MTP predicts the next tokens a…
cs.LG2025
Don't Mesh with Me: Generating Constructive Solid Geometry Instead of Meshes by Fine-Tuning a Code-Generation LLM
Maximilian Mews, Ansar Aynetdinov, Vivian Schiller +2
While recent advancements in machine learning, such as LLMs, are revolutionizing software development and creative industries, they have had minimal impact on engineers designing m…