2 papers
cs.LG2026
Breaking the Illusion: When Positive Meets Negative in Multimodal Decoding
Yubo Jiang, Yitong An, Xin Yang +7
Vision-Language Models (VLMs) are frequently undermined by object hallucination, generating content that contradicts visual reality, due to an over-reliance on linguistic priors. W…
cs.AI2026
V-tableR1: Process-Supervised Multimodal Table Reasoning with Critic-Guided Policy Optimization
Yubo Jiang, Yitong An, Xin Yang +7
We introduce V-tableR1, a process-supervised reinforcement learning framework that elicits rigorous, verifiable reasoning from multimodal large language models (MLLMs). Current MLL…