8 papers
Test-Time Training with Next-Token Prediction
Xuan Ouyang, Zefan Cai, Junjie Hu
Next-token prediction is the self-supervised signal that trains language models, and every observed prompt token provides the same signal at test time. We study whether this signal…
EvoPool: Evolutionary Programmatic Annotation for Label-Efficient Specialized Supervision
Tianyi Xu, Yaolun Zhang, Xuan Ouyang +1
Large language models excel at general tasks but underperform smaller supervised models in specialized, high-stakes domains where training labels are costly. We address this regime…
TRACE: Distilling Where It Matters via Token-Routed Self On-Policy Alignment
Jiaxuan Wang, Xuan Ouyang, Zhiyu Chen +4
On-policy self-distillation (self-OPD) densifies reinforcement learning with verifiable rewards (RLVR) by letting a policy teach itself under privileged context. We find that when…
LOLGORITHM: Funny Comment Generation Agent For Short Videos
Xuan Ouyang, Bouzhou Wang, Senan Wang +3
Short-form video platforms have become central to multimedia information dissemination, where comments play a critical role in driving engagement, propagation, and algorithmic feed…
The Side Effects of Being Smart: Safety Risks in MLLMs' Multi-Image Reasoning
Renmiao Chen, Yida Lu, Shiyao Cui +6
As Multimodal Large Language Models (MLLMs) acquire stronger reasoning capabilities to handle complex, multi-image instructions, this advancement may pose new safety risks. We stud…
Laugh, Relate, Engage: Stylized Comment Generation for Short Videos
Xuan Ouyang, Senan Wang, Bouzhou Wang +3
Short-video platforms have become a central medium in the modern Internet landscape, where efficient information delivery and strong interactivity are reshaping user engagement and…