4 citations · 5 across the 2 of their papers we have counts for
4 papers · 1 filter
Mettle: Meta-Token Learning for Memory-Efficient Audio-Visual Adaptation
Jinxing Zhou, Zhihui Li, Yongqiang Yu +7
We present \textbf{Met}a-\textbf{T}oken \textbf{Le}arning (Mettle), a simple and memory-efficient method for adapting large-scale pretrained transformer models to downstream audio-…
StoryAgent: Customized Storytelling Video Generation via Multi-Agent Collaboration
Panwen Hu, Jin Jiang, Jianqi Chen +4
The advent of AI-Generated Content (AIGC) has spurred research into automated video generation to streamline conventional processes. However, automating storytelling video producti…
Mask Propagation for Efficient Video Semantic Segmentation
Yuetian Weng, Mingfei Han, Haoyu He +4
Video Semantic Segmentation (VSS) involves assigning a semantic label to each pixel in a video sequence. Prior work in this field has demonstrated promising results by extending im…
An Efficient Spatio-Temporal Pyramid Transformer for Action Detection
Yuetian Weng, Zizheng Pan, Mingfei Han +2
The task of action detection aims at deducing both the action category and localization of the start and end moment for each action instance in a long, untrimmed video. While visio…