collaborators

5 papers

cs.LG2026

Data-Constrained Language Model Pretraining: Improved Regularization and Scaling Laws

Zhiwei Xu, Shihao Wu, Hanseul Cho +2

Classical scaling laws for language model pretraining balance model size against training dataset size under a fixed compute budget, assuming abundant data and a single pass over t…

cs.SE2026

Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Mike A. Merrill, Alexander G. Shaw, Nicholas Carlini +82

AI agents may soon become capable of autonomously completing valuable, long-horizon tasks in diverse domains. Current benchmarks either do not measure real-world tasks, or are not…

stat.ME2025

Inference of Dependency Knowledge Graph for Electronic Health Records

Zhiwei Xu, Ziming Gan, Doudou Zhou +3

The effective analysis of high-dimensional Electronic Health Record (EHR) data, with substantial potential for healthcare research, presents notable methodological challenges. Empl…

cs.LG2025

Let Me Grok for You: Accelerating Grokking via Embedding Transfer from a Weaker Model

Zhiwei Xu, Zhiyu Ni, Yixin Wang +1

''Grokking'' is a phenomenon where a neural network first memorizes training data and generalizes poorly, but then suddenly transitions to near-perfect generalization after prolong…

cs.LG2025

Benign Overfitting in Single-Head Attention

Roey Magen, Shuning Shang, Zhiwei Xu +3

The phenomenon of benign overfitting, where a trained neural network perfectly fits noisy training data but still achieves near-optimal test performance, has been extensively studi…