collaborators

8 papers

cs.CL2026

Test-Time Training with Next-Token Prediction

Xuan Ouyang, Zefan Cai, Junjie Hu

Next-token prediction is the self-supervised signal that trains language models, and every observed prompt token provides the same signal at test time. We study whether this signal…

cs.CL2026

EvoPool: Evolutionary Programmatic Annotation for Label-Efficient Specialized Supervision

Tianyi Xu, Yaolun Zhang, Xuan Ouyang +1

Large language models excel at general tasks but underperform smaller supervised models in specialized, high-stakes domains where training labels are costly. We address this regime…

cs.AI2026

TRACE: Distilling Where It Matters via Token-Routed Self On-Policy Alignment

Jiaxuan Wang, Xuan Ouyang, Zhiyu Chen +4

On-policy self-distillation (self-OPD) densifies reinforcement learning with verifiable rewards (RLVR) by letting a policy teach itself under privileged context. We find that when…

cs.CV2026

LOLGORITHM: Funny Comment Generation Agent For Short Videos

Xuan Ouyang, Bouzhou Wang, Senan Wang +3

Short-form video platforms have become central to multimedia information dissemination, where comments play a critical role in driving engagement, propagation, and algorithmic feed…

cs.CV2026

The Side Effects of Being Smart: Safety Risks in MLLMs' Multi-Image Reasoning

Renmiao Chen, Yida Lu, Shiyao Cui +6

As Multimodal Large Language Models (MLLMs) acquire stronger reasoning capabilities to handle complex, multi-image instructions, this advancement may pose new safety risks. We stud…

cs.LG2025

Laugh, Relate, Engage: Stylized Comment Generation for Short Videos

Xuan Ouyang, Senan Wang, Bouzhou Wang +3

Short-video platforms have become a central medium in the modern Internet landscape, where efficient information delivery and strong interactivity are reshaping user engagement and…