3 papers
cs.CV2024
Grounding is All You Need? Dual Temporal Grounding for Video Dialog
You Qin, Wei Ji, Xinze Lan +5
In the realm of video dialog response generation, the understanding of video content and the temporal nuances of conversation history are paramount. While a segment of current rese…
cs.CV2024
DriveDiTFit: Fine-tuning Diffusion Transformers for Autonomous Driving
Jiahang Tu, Wei Ji, Hanbin Zhao +3
In autonomous driving, deep models have shown remarkable performance across various visual perception tasks with the demand of high-quality and huge-diversity training datasets. Su…
cs.CV2024
Described Spatial-Temporal Video Detection
Wei Ji, Xiangyan Liu, Yingfei Sun +6
Detecting visual content on language expression has become an emerging topic in the community. However, in the video domain, the existing setting, i.e., spatial-temporal video grou…