2 papers
cs.CL2026
TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs
Andong Hua, Colton Bishop, Igor Mordatch +5
Multimodal large language models (MLLMs) should generate consistent responses given semantically equivalent inputs across modalities. However, we observe a systematic discrepancy i…
cs.LG2024
Training Language Models to Self-Correct via Reinforcement Learning
Aviral Kumar, Vincent Zhuang, Rishabh Agarwal +15
Self-correction is a highly desirable capability of large language models (LLMs), yet it has consistently been found to be largely ineffective in modern LLMs. Current methods for t…