collaborators

9 papers

cs.IR2026

Co-Scraper: query-aware DOM Pruning and Reusable Scraper Synthesis for Lightweight Web Data Extraction

Shoupeng Wang, Jiantao Qiu, Wuyang Zhang +1

The abundant and heterogeneous nature of web content necessitates automated information extraction, and generating scrapers that can be reused across similar web pages offers an ef…

cs.LG2026

Intern-S1-Pro: Scientific Multimodal Foundation Model at Trillion Scale

Yicheng Zou, Dongsheng Zhu, Lin Zhu +174

We introduce Intern-S1-Pro, the first one-trillion-parameter scientific multimodal foundation model. Scaling to this unprecedented size, the model delivers a comprehensive enhancem…

physics.chem-ph2026

NMRTrans: Structure Elucidation from Experimental NMR Spectra via Set Transformers

Liujia Yang, Zhuo Yang, Jiaqing Xie +9

Nuclear Magnetic Resonance (NMR) spectroscopy is fundamental for molecular structure elucidation, yet interpreting spectra at scale remains time-consuming and highly expertise-depe…

cs.CL2025

Topic Over Source: The Key to Effective Data Mixing for Language Models Pre-training

Jiahui Peng, Xinlin Zhuang, Jiantao Qiu +4

The performance of large language models (LLMs) is significantly affected by the quality and composition of their pre-training data, which is inherently diverse, spanning various l…

cs.CL2025

Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models

Xinlin Zhuang, Jiahui Peng, Ren Ma +7

The composition of pre-training datasets for large language models (LLMs) remains largely undisclosed, hindering transparency and efforts to optimize data quality, a critical drive…

cs.CL2025

Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration

Tianyi Bai, Ling Yang, Zhen Hao Wong +9

Efficient data selection is crucial to accelerate the pretraining of language model (LMs). While various methods have been proposed to enhance data efficiency, limited research has…