15 citations · 19 across the 6 of their papers we have counts for
8 papers
MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation
Xiaohan Zhang, Yuqing Wen, Junlin Chen +9
Multi-reference-to-audio-video (MR2AV) generation aims to generate coherent audio-video content conditioned on multiple references and textual instructions. Existing benchmarks mai…
Scalable In-Context Q-Learning
Jinmei Liu, Fuhong Liu, Zhenhong Sun +6
Recent advancements in language models have demonstrated remarkable in-context learning abilities, prompting the exploration of in-context reinforcement learning (ICRL) to extend t…
Fast Disentangled Slim Tensor Learning for Multi-view Clustering
Deng Xu, Chao Zhang, Zechao Li +2
Tensor-based multi-view clustering has recently received significant attention due to its exceptional ability to explore cross-view high-order correlations. However, most existing…
DeMamba: AI-Generated Video Detection on Million-Scale GenVideo Benchmark
Haoxing Chen, Yan Hong, Zizheng Huang +8
Recently, video generation techniques have advanced rapidly. Given the popularity of video content on social media platforms, these models intensify concerns about the spread of fa…
Segment Anything Model Meets Image Harmonization
Haoxing Chen, Yaohui Li, Zhangxuan Gu +3
Image harmonization is a crucial technique in image composition that aims to seamlessly match the background by adjusting the foreground of composite images. Current methods adopt…
Contrastive Latent Space Reconstruction Learning for Audio-Text Retrieval
Kaiyi Luo, Xulong Zhang, Jianzong Wang +3
Cross-modal retrieval (CMR) has been extensively applied in various domains, such as multimedia search engines and recommendation systems. Most existing CMR methods focus on image-…