4 papers
LLM4VG: Large Language Models Evaluation for Video Grounding
Wei Feng, Xin Wang, Hong Chen +7
Recently, researchers have attempted to investigate the capability of LLMs in handling videos and proposed several video LLM models. However, the ability of LLMs to handle video gr…
Multi-sentence Video Grounding for Long Video Generation
Wei Feng, Xin Wang, Hong Chen +2
Video generation has witnessed great success recently, but their application in generating long videos still remains challenging due to the difficulty in maintaining the temporal c…
RealTCD: Temporal Causal Discovery from Interventional Data with Large Language Model
Peiwen Li, Xin Wang, Zeyang Zhang +6
In the field of Artificial Intelligence for Information Technology Operations, causal discovery is pivotal for operation and maintenance of graph construction, facilitating downstr…
DisenStudio: Customized Multi-subject Text-to-Video Generation with Disentangled Spatial Control
Hong Chen, Xin Wang, Yipeng Zhang +4
Generating customized content in videos has received increasing attention recently. However, existing works primarily focus on customized text-to-video generation for single subjec…