4 papers
AHPA: Adaptive Hierarchical Prior Alignment for Diffusion Transformers
Ruibin Min, Yexin Liu, Aimin Pan +5
Representation alignment has recently emerged as an effective paradigm for accelerating Diffusion Transformer training. Despite their success, existing alignment methods typically…
RemoteShield: Enable Robust Multimodal Large Language Models for Earth Observation
Rui Min, Liang Yao, Shiyu Miao +5
A robust Multimodal Large Language Model (MLLM) for Earth Observation should maintain consistent interpretation and reasoning under realistic input variations. However, current Rem…
PairJudge RM: Perform Best-of-N Sampling with Knockout Tournament
Yantao Liu, Zijun Yao, Rui Min +3
Best-of-N (BoN) sampling, a common strategy for test-time scaling of Large Language Models (LLMs), relies on reward models to select the best candidate solution from multiple gener…
RM-Bench: Benchmarking Reward Models of Language Models with Subtlety and Style
Yantao Liu, Zijun Yao, Rui Min +3
Reward models are critical in techniques like Reinforcement Learning from Human Feedback (RLHF) and Inference Scaling Laws, where they guide language model alignment and select opt…