biomedical text generation 1medication leaflets 1post-training alignment 1preference optimization 1small language models 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.CL2026
Exploring Post-Training Alignment of Small Language Models for Biomedical Data-to-Text Generation: A Case Study of Medication Leaflet
Xi Yang, Guodong Liu, Chuqin Li +10
The paper compares several post‑training alignment methods for small language models on the task of converting biomedical data into patient‑friendly medication leaflets, showing th…
cs.LG2026
One Student, Many Teachers: Multi-Task On-Policy Distillation via Soft-Prompt Privileged Context
Yingzi Ma, Zichen Zhu, Ming Jiang +1
On-policy self-distillation (OPSD) teaches large language models new skills through a teacher that shares the student's backbone and supervises its own rollouts. Existing teachers…