2 papers
cs.LG2026
RAIN-Merging: A Gradient-Free Method to Enhance Instruction Following in Large Reasoning Models with Preserved Thinking Format
Zhehao Huang, Yuhang Liu, Baijiong Lin +5
Large reasoning models (LRMs) excel at a long chain of reasoning but often fail to faithfully follow instructions regarding output format, constraints, or specific requirements. We…
cs.CV2025
T2I-ConBench: Text-to-Image Benchmark for Continual Post-training
Zhehao Huang, Yuhang Liu, Yixin Lou +7
Continual post-training adapts a single text-to-image diffusion model to learn new tasks without incurring the cost of separate models, but naive post-training causes forgetting of…