64 citations · 64 across the 3 of their papers we have counts for
3 papers
cs.CV2025
When and Where do Events Switch in Multi-Event Video Generation?
Ruotong Liao, Guowen Huang, Qing Cheng +3
Text-to-video (T2V) generation has surged in response to challenging questions, especially when a long video must depict multiple sequential events with temporal coherence and cont…
cs.CL2023
GraphextQA: A Benchmark for Evaluating Graph-Enhanced Large Language Models
Yuanchun Shen, Ruotong Liao, Zhen Han +2
While multi-modal models have successfully integrated information from image, video, and audio modalities, integrating graph modality into large language models (LLMs) remains unex…
cs.CV2023★ 64 cited
A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models
Jindong Gu, Zhen Han, Shuo Chen +7
Prompt engineering is a technique that involves augmenting a large pre-trained model with task-specific hints, known as prompts, to adapt the model to new tasks. Prompts can be cre…