1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.CV2024★ 1 cited
MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos
Xuehai He, Weixi Feng, Kaizhi Zheng +11
Multimodal Language Language Models (MLLMs) demonstrate the emerging abilities of "world models" -- interpreting and reasoning about complex real-world dynamics. To assess these ab…
cs.CV2023
Multimodal Graph Transformer for Multimodal Question Answering
Xuehai He, Xin Eric Wang
Despite the success of Transformer models in vision and language tasks, they often learn knowledge from enormous data implicitly and cannot utilize structured input data directly.…