3 papers
cs.LG2026
Two to Tango: Coupled Task-Reference Selection for Safe LLM Fine-tuning
Xinrui Chen, Jianhao Zhang, Ou Wu +1
Fine-tuning safety aligned large language models (LLMs) on downstream data improves adaptation but may erode learned safety behavior. Existing methods use fixed safety examples, gl…
cs.CL2025
Skywork-R1V3 Technical Report
Wei Shen, Jiangbo Pei, Yi Peng +8
We introduce Skywork-R1V3, an advanced, open-source vision-language model (VLM) that pioneers a new approach to visual reasoning. Its key innovation lies in effectively transferrin…
cs.CV2025
Skywork R1V2: Multimodal Hybrid Reinforcement Learning for Reasoning
Peiyu Wang, Yichen Wei, Yi Peng +10
We present Skywork R1V2, a next-generation multimodal reasoning model and a major leap forward from its predecessor, Skywork R1V. At its core, R1V2 introduces a hybrid reinforcemen…