8 papers
On Distinguishing Capability Elicitation from Capability Creation in Post-Training: A Free-Energy Perspective
Yuhao Li, Shengchao Liu
Debates about large language model post-training often treat supervised fine-tuning (SFT) as imitation and reinforcement learning (RL) as discovery. But this distinction is too coa…
HER: Human-like Reasoning and Reinforcement Learning for LLM Role-playing
Chengyu Du, Xintao Wang, Aili Chen +11
LLM role-playing, i.e., using LLMs to simulate specific personas, has emerged as a key capability in various applications, such as companionship, content creation and digital games…
A Minimal Model of Representation Collapse: Frustration, Stop-Gradient, and Dynamics
Louie Hong Yao, Yuhao Li, Shengchao Liu
Self-supervised representation learning is central to modern machine learning because it extracts structured latent features from unlabeled data and enables robust transfer across…
Finding Bugs in Short Proofs: The Metamathematics of Resolution Lower Bounds
Jiawei Li, Yuhao Li, Hanlin Ren
We study the *refuter* problems for proof complexity lower bounds. Suppose is a hard tautology that does not admit any length- proof in some proof system . In the corres…
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
MiniMax, :, Aili Chen +125
We introduce MiniMax-M1, the world's first open-weight, large-scale hybrid-attention reasoning model. MiniMax-M1 is powered by a hybrid Mixture-of-Experts (MoE) architecture combin…
Dynamical Label Augmentation and Calibration for Noisy Electronic Health Records
Yuhao Li, Ling Luo, Uwe Aickelin
Medical research, particularly in predicting patient outcomes, heavily relies on medical time series data extracted from Electronic Health Records (EHR), which provide extensive in…