collaborators

9 papers

physics.flu-dyn2026

DD-RNO: A Domain-Decomposed Routed Neural Operator for Airfoil Flow Prediction

T. A. Mehta, P. S. Bhati, H. D. Akolekar

Deep learning surrogates for RANS flow prediction around airfoils face two persistent bottlenecks. A single neural architecture cannot simultaneously resolve sharp near-wall bounda…

cs.CL2026

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data

Xu Guo, Runyu Peng, Jian Tong +4

Large language models (LLMs) rely on web-scale corpora for pre-training. The noise inherent in these datasets tends to obscure meaningful patterns and ultimately degrade model perf…

cs.LG2026

Intern-S1-Pro: Scientific Multimodal Foundation Model at Trillion Scale

Yicheng Zou, Dongsheng Zhu, Lin Zhu +174

We introduce Intern-S1-Pro, the first one-trillion-parameter scientific multimodal foundation model. Scaling to this unprecedented size, the model delivers a comprehensive enhancem…

cs.LG2026

What Makes Position Zero Special? A Mechanistic Study of Position Zero Attention Sinks in LLMs

Runyu Peng, Ruixiao Li, Mingshu Chen +5

Transformers frequently allocate disproportionate attention to specific tokens, a phenomenon known as attention sinks. Causal large language models reliably form one at position ze…

cs.LG2026

Explicit Multi-head Attention for Inter-head Interaction in Large Language Models

Runyu Peng, Yunhua Zhou, Demin Song +4

In large language models built upon the Transformer architecture, recent studies have shown that inter-head interaction can enhance attention performance. Motivated by this, we pro…

cs.AI2026

How to Set the Batch Size for Large-Scale Pre-training?

Yunhua Zhou, Junhao Huang, Shuhao Xing +4

The concept of Critical Batch Size, as pioneered by OpenAI, has long served as a foundational principle for large-scale pre-training. However, with the paradigm shift towards the W…