11 citations · 23 across the 3 of their papers we have counts for
3 papers
cs.CV2024★ 3 cited
ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Lin Chen, Xilin Wei, Jinsong Li +12
We present the ShareGPT4Video series, aiming to facilitate the video understanding of large video-language models (LVLMs) and the video generation of text-to-video models (T2VMs) v…
cs.CV2024★ 11 cited
InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model
Xiaoyi Dong, Pan Zhang, Yuhang Zang +20
We introduce InternLM-XComposer2, a cutting-edge vision-language model excelling in free-form text-image composition and comprehension. This model goes beyond conventional vision-l…
cs.CL2023★ 9 cited
Evaluating and Improving Tool-Augmented Computation-Intensive Math Reasoning
Beichen Zhang, Kun Zhou, Xilin Wei +4
Chain-of-thought prompting~(CoT) and tool augmentation have been validated in recent work as effective practices for improving large language models~(LLMs) to perform step-by-step…