71 citations · 108 across the 8 of their papers we have counts for
9 papers
MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos
Xuehai He, Weixi Feng, Kaizhi Zheng +11
Multimodal Language Language Models (MLLMs) demonstrate the emerging abilities of "world models" -- interpreting and reasoning about complex real-world dynamics. To assess these ab…
MMSci: A Dataset for Graduate-Level Multi-Discipline Multimodal Scientific Understanding
Zekun Li, Xianjun Yang, Kyuri Choi +11
Scientific figure interpretation is a crucial capability for AI-driven scientific assistants built on advanced Large Vision Language Models. However, current datasets and benchmark…
Automatic Layout Planning for Visually-Rich Documents with Instruction-Following Models
Wanrong Zhu, Jennifer Healey, Ruiyi Zhang +2
Recent advancements in instruction-following models have made user interactions with models more user-friendly and efficient, broadening their applicability. In graphic design, non…
CLIP also Understands Text: Prompting CLIP for Phrase Understanding
An Yan, Jiacheng Li, Wanrong Zhu +3
Contrastive Language-Image Pretraining (CLIP) efficiently learns visual concepts by pre-training with natural language supervision. CLIP and its visual encoder have been explored o…
Imagination-Augmented Natural Language Understanding
Yujie Lu, Wanrong Zhu, Xin Eric Wang +2
Human brains integrate linguistic and perceptual information simultaneously to understand natural language, and hold the critical ability to render imaginations. Such abilities ena…
Towards Understanding Sample Variance in Visually Grounded Language Generation: Evaluations and Observations
Wanrong Zhu, Xin Eric Wang, Pradyumna Narayana +3
A major challenge in visually grounded language generation is to build robust benchmark datasets and models that can generalize well in real-world settings. To do this, it is criti…