From the 1 of 6 linked papers with an AI index.
6 papers
RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination
Haotian Liang, Mingkang Chen, Yufei Huang +27
The paper introduces RxBrain, a foundation model that jointly reasons over language and visual inputs to create embodied plans, using a multimodal Mixture-of-Transformers architect…
Global Context or Local Detail? Adaptive Visual Grounding for Hallucination Mitigation
Yubo Jiang, Xin Yang, Abudukelimu Wuerkaixi +7
Vision-Language Models (VLMs) are frequently undermined by object hallucination--generating content that contradicts visual reality--due to an over-reliance on linguistic priors. W…
Breaking the Illusion: When Positive Meets Negative in Multimodal Decoding
Yubo Jiang, Yitong An, Xin Yang +7
Vision-Language Models (VLMs) are frequently undermined by object hallucination, generating content that contradicts visual reality, due to an over-reliance on linguistic priors. W…
V-tableR1: Process-Supervised Multimodal Table Reasoning with Critic-Guided Policy Optimization
Yubo Jiang, Yitong An, Xin Yang +7
We introduce V-tableR1, a process-supervised reinforcement learning framework that elicits rigorous, verifiable reasoning from multimodal large language models (MLLMs). Current MLL…
CoPA: Hierarchical Concept Prompting and Aggregating Network for Explainable Diagnosis
Yiheng Dong, Yi Lin, Xin Yang
The transparency of deep learning models is essential for clinical diagnostics. Concept Bottleneck Model provides clear decision-making processes for diagnosis by transforming the…
MANBench: Is Your Multimodal Model Smarter than Human?
Han Zhou, Qitong Xu, Yiheng Dong +1
The rapid advancement of Multimodal Large Language Models (MLLMs) has ignited discussions regarding their potential to surpass human performance in multimodal tasks. In response, w…