2 papers
cs.CV2022
Learning Fine-Grained Visual Understanding for Video Question Answering via Decoupling Spatial-Temporal Modeling
Hsin-Ying Lee, Hung-Ting Su, Bing-Chen Tsai +3
While recent large-scale video-language pre-training made great progress in video question answering, the design of spatial modeling of video-language models is less fine-grained t…
cs.AI2021
TrUMAn: Trope Understanding in Movies and Animations
Hung-Ting Su, Po-Wei Shen, Bing-Chen Tsai +3
Understanding and comprehending video content is crucial for many real-world applications such as search and recommendation systems. While recent progress of deep learning has boos…