5 papers
Data-Constrained Language Model Pretraining: Improved Regularization and Scaling Laws
Zhiwei Xu, Shihao Wu, Hanseul Cho +2
Classical scaling laws for language model pretraining balance model size against training dataset size under a fixed compute budget, assuming abundant data and a single pass over t…
Efficient Synthetic Network Generation via Latent Embedding Reconstruction
Feifan Jiang, Yinan Bu, Shihao Wu +2
Network data are ubiquitous across the social sciences, biology, and information systems. Generating realistic synthetic network data has broad applications from network simulation…
Denoising Diffused Embeddings: a Generative Approach for Hypergraphs
Shihao Wu, Junyi Yang, Gongjun Xu +1
Hypergraph data, which capture multi-way interactions among entities, are increasingly prevalent in the big data era. Generating new hyperlinks from an observed, usually high-dimen…
A General Latent Embedding Approach for Modeling Non-uniform High-dimensional Sparse Hypergraphs with Multiplicity
Shihao Wu, Gongjun Xu, Ji Zhu
Recent research has shown growing interest in modeling hypergraphs, which capture polyadic interactions among entities beyond traditional dyadic relations. However, most existing m…
Statistical Inference on Latent Space Models for Network Data
Jinming Li, Shihao Wu, Chengyu Cui +2
Latent space models are powerful statistical tools for modeling and understanding network data. While the importance of accounting for uncertainty in network analysis has been well…