1 citations · 1 across the 3 of their papers we have counts for
4 papers
Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance
Yang Zhang, Chenwei Wang, Ouyang Lu +7
Vision-Language-Action (VLA) models pre-trained on large, diverse datasets show remarkable potential for general-purpose robotic manipulation. However, a primary bottleneck remains…
Online Iterative Self-Alignment for Radiology Report Generation
Ting Xiao, Lei Shi, Yang Zhang +3
Radiology Report Generation (RRG) is an important research topic for relieving radiologist' heavy workload. Existing RRG models mainly rely on supervised fine-tuning (SFT) based on…
Revisiting Multi-Agent World Modeling from a Diffusion-Inspired Perspective
Yang Zhang, Xinran Li, Jianing Ye +5
World models have recently attracted growing interest in Multi-Agent Reinforcement Learning (MARL) due to their ability to improve sample efficiency for policy learning. However, a…
Online Preference Alignment for Language Models via Count-based Exploration
Chenjia Bai, Yang Zhang, Shuang Qiu +3
Reinforcement Learning from Human Feedback (RLHF) has shown great potential in fine-tuning Large Language Models (LLMs) to align with human preferences. Existing methods perform pr…