1 paper
Lei Zhao, Zihao Ma, Boyu Lin +3
We present an RL-central framework for Language and Vision Assistants (RLLaVA) with its formulation of Markov decision process (MDP). RLLaVA decouples RL algorithmic logic from mod…