3 citations · 3 across the 3 of their papers we have counts for
1 paper · 1 filter
Nikhil Verma, Minjung Kim, JooYoung Yoo +5
Multimodal language models now integrate text, audio, and video for unified reasoning. Yet existing RL post-training pipelines treat all input signals as equally relevant, ignoring…