2 papers
cs.SD2025
TAIL: Text-Audio Incremental Learning
Yingfei Sun, Xu Gu, Wei Ji +3
Many studies combine text and audio to capture multi-modal information but they overlook the model's generalization ability on new datasets. Introducing new datasets may affect the…
cs.CV2024
Described Spatial-Temporal Video Detection
Wei Ji, Xiangyan Liu, Yingfei Sun +6
Detecting visual content on language expression has become an emerging topic in the community. However, in the video domain, the existing setting, i.e., spatial-temporal video grou…