1 citations · 2 across the 8 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2023★ 1 cited
Video-Teller: Enhancing Cross-Modal Generation with Fusion and Decoupling
Haogeng Liu, Qihang Fan, Tingkai Liu +5
This paper proposes Video-Teller, a video-language foundation model that leverages multi-modal fusion and fine-grained modality alignment to significantly enhance the video-to-text…
cs.CV2023
DeVAn: Dense Video Annotation for Video-Language Models
Tingkai Liu, Yunzhe Tao, Haogeng Liu +5
We present a novel human annotated dataset for evaluating the ability for visual-language models to generate both short and long descriptions for real-world video clips, termed DeV…
cs.CV2023
Expedited Training of Visual Conditioned Language Generation via Redundancy Reduction
Yiren Jian, Tingkai Liu, Yunzhe Tao +3
In this paper, we introduce , a streamlined framework designed for the pre-training of visually conditioned language generation models with high computatio…