6 papers
Mitigating Identity Essentialism in LLM Agents with Longitudinal Life Trajectories
Hexi Wang, Yujia Zhou, Bangde Du +7
Large language models (LLMs) offer a scalable approach to social simulation, but their credibility depends on how agents are constructed. Existing methods can partially reproduce p…
Retrievable Gradients: Continual Post-Training Without Cumulative Weight Drift
Weihang Su, Jiacheng Kang, Jingyan Xu +7
Continual post-training enables models to absorb emerging knowledge after deployment, but repeatedly updating shared parameters can accumulate weight drift, potentially causing cat…
Continuous Latent Contexts Enable Efficient Online Learning in Transformers
Emile Anand, Abdullah Ateyeh, Xinyuan Cao +1
Large language models (LLMs) exhibit a strong capacity for in-context learning: Given labeled examples, they can generate good predictions without parameter updates. However, many…
Provable Long-Range Benefits of Next-Token Prediction
Xinyuan Cao, Santosh S. Vempala
Why do modern language models, trained to do well on next-word prediction, appear to generate coherent documents and capture long-range structure? Here we show that next-token pred…
StructComp: Substituting Propagation with Structural Compression in Training Graph Contrastive Learning
Shengzhong Zhang, Wenjie Yang, Xinyuan Cao +2
Graph contrastive learning (GCL) has become a powerful tool for learning graph data, but its scalability remains a significant challenge. In this work, we propose a simple yet effe…
Contrastive Moments: Unsupervised Halfspace Learning in Polynomial Time
Xinyuan Cao, Santosh S. Vempala
We give a polynomial-time algorithm for learning high-dimensional halfspaces with margins in -dimensional space to within desired TV distance when the ambient distribution is an…