diffusion models 1language model adaptation 1screenwriting generation 1self-supervised learning 1skill extraction 1
From the 1 of 12 linked papers with an AI index.
3 citations · 3 across the 7 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026★ 3 cited
VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models
Haodong Duan, Xinyu Fang, Junming Yang +41
We present VLMEvalKit: an open-source toolkit for evaluating large multi-modality models based on PyTorch. The toolkit aims to provide a user-friendly and comprehensive framework f…
cs.CV2026
MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence
Sihan Yang, Runsen Xu, Yiman Xie +10
Spatial intelligence is essential for multimodal large language models (MLLMs) operating in the complex physical world. Existing benchmarks, however, probe only single-image relati…