1 paper
Jianghao Wu, Yasmeen George, Jin Ye +3
Large language models (LLMs) and multimodal LLMs (MLL-Ms) excel at chain-of-thought reasoning but face distribution shift at test-time and a lack of verifiable supervision. Recent…