3 papers
cs.LG2026
Twin Rollouts: Noise-Coupled Counterfactual Branching in Interactive Video World Models
Yu Ma, Hongli Shi, Xinran Xu
Interactive video world models generate rollouts autoregressively under an action stream, yet they are trained and evaluated almost exclusively on factual prediction. We study coun…
cs.LG2026
When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs
Yu Ma, Hongli Shi, Jing Li +2
Model merging has become the default way to give an aligned language model new skills without retraining: a practitioner folds task vectors from math, code, or domain specialists i…
cs.SE2024
Unlock the Correlation between Supervised Fine-Tuning and Reinforcement Learning in Training Code Large Language Models
Jie Chen, Xintian Han, Yu Ma +2
Automatic code generation has been a longstanding research topic. With the advancement of general-purpose large language models (LLMs), the ability to code stands out as one import…