4 papers · 1 filter
OSCBench: Benchmarking Object State Change in Text-to-Video Generation
Xianjing Han, Bin Zhu, Shiqi Hu +4
Text-to-video (T2V) generation models have made rapid progress in producing visually high-quality and temporally coherent videos. However, existing benchmarks primarily focus on pe…
Grounding is All You Need? Dual Temporal Grounding for Video Dialog
You Qin, Wei Ji, Xinze Lan +5
In the realm of video dialog response generation, the understanding of video content and the temporal nuances of conversation history are paramount. While a segment of current rese…
DriveDiTFit: Fine-tuning Diffusion Transformers for Autonomous Driving
Jiahang Tu, Wei Ji, Hanbin Zhao +3
In autonomous driving, deep models have shown remarkable performance across various visual perception tasks with the demand of high-quality and huge-diversity training datasets. Su…
Described Spatial-Temporal Video Detection
Wei Ji, Xiangyan Liu, Yingfei Sun +6
Detecting visual content on language expression has become an emerging topic in the community. However, in the video domain, the existing setting, i.e., spatial-temporal video grou…