2 papers
cs.CL2026
UltraX: Refining Pre-Training Data at Scale with Adaptive Programmatic Editing
Xinlong Zhao, Dongsheng Liu, Hengyu Zhao +9
As available training data approaches its physical limit, gains from Scaling Laws have begun to diminish. Consequently, improving Large Language Models (LLMs) now depends less on d…
cs.PF2026
When Is the Same Model Not the Same Service? A Measurement Study of Hosted Open-Weight LLM APIs
Haorui Li, Zhenghui He, Xuanzi Liu +7
Open-weight large language models (LLMs) are usually named as model artifacts, but production users often consume them as hosted API services. This paper argues that the operationa…