2 papers
cs.LG2026
OISD: On-Policy Internal Self-Distillation of Language Models
Xinyu Liu, Darryl Cherian Jacob, Yang Zhou +2
Recent reinforcement learning (RL) post-training approaches primarily optimize the final output policy using sparse outcome-level rewards, while largely overlooking predictive sign…
cs.AI2026
TorchUMM: A Unified Multimodal Model Codebase for Evaluation, Analysis, and Post-training
Yinyi Luo, Wenwen Wang, Hayes Bai +6
Recent advances in unified multimodal models (UMMs) have led to a proliferation of architectures capable of understanding, generating, and editing across visual and textual modalit…