8 citations · 19 across the 6 of their papers we have counts for
6 papers
Computing in the Era of Large Generative Models: From Cloud-Native to AI-Native
Yao Lu, Song Bian, Lequn Chen +19
In this paper, we investigate the intersection of large generative AI models and cloud-native computing architectures. Recent large models such as ChatGPT, while revolutionary in t…
Large Model based Sequential Keyframe Extraction for Video Summarization
Kailong Tan, Yuxiang Zhou, Qianchen Xia +2
Keyframe extraction aims to sum up a video's semantics with the minimum number of its frames. This paper puts forward a Large Model based Sequential Keyframe Extraction for video s…
Bird's-Eye-View Scene Graph for Vision-Language Navigation
Rui Liu, Xiaohan Wang, Wenguan Wang +1
Vision-language navigation (VLN), which entails an agent to navigate 3D environments following human instructions, has shown great advances. However, current agents are built upon…
Betray Oneself: A Novel Audio DeepFake Detection Model via Mono-to-Stereo Conversion
Rui Liu, Jinhua Zhang, Guanglai Gao +1
Audio Deepfake Detection (ADD) aims to detect the fake audio generated by text-to-speech (TTS), voice conversion (VC) and replay, etc., which is an emerging topic. Traditionally we…
PerCoNet: News Recommendation with Explicit Persona and Contrastive Learning
Rui Liu, Bin Yin, Ziyi Cao +3
Personalized news recommender systems help users quickly find content of their interests from the sea of information. Today, the mainstream technology for personalized news recomme…
PP-MSVSR: Multi-Stage Video Super-Resolution
Lielin Jiang, Na Wang, Qingqing Dang +2
Different from the Single Image Super-Resolution(SISR) task, the key for Video Super-Resolution(VSR) task is to make full use of complementary information across frames to reconstr…