222 citations · 534 across the 41 of their papers we have counts for
16 papers · 1 filter
SIGMark: Scalable In-Generation Watermark with Blind Extraction for Video Diffusion
Xinjie Zhu, Zijing Zhao, Hui Jin +5
Artificial Intelligence Generated Content (AIGC), particularly video generation with diffusion models, has been advanced rapidly. Invisible watermarking is a key technology for pro…
MRN: Multiplexed Routing Network for Incremental Multilingual Text Recognition
Tianlun Zheng, Zhineng Chen, BingChen Huang +2
Multilingual text recognition (MLTR) systems typically focus on a fixed set of languages, which makes it difficult to handle newly added languages or adapt to ever-changing data di…
LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model
Peng Gao, Jiaming Han, Renrui Zhang +9
How to efficiently transform large language models (LLMs) into instruction followers is recently a popular research direction, while training LLM for multi-modal reasoning remains…
Network Pruning Spaces
Xuanyu He, Yu-I Yang, Ran Song +5
Network pruning techniques, including weight pruning and filter pruning, reveal that most state-of-the-art neural networks can be accelerated without a significant performance drop…
DetCLIPv2: Scalable Open-Vocabulary Object Detection Pre-training via Word-Region Alignment
Lewei Yao, Jianhua Han, Xiaodan Liang +4
This paper presents DetCLIPv2, an efficient and scalable training framework that incorporates large-scale image-text pairs to achieve open-vocabulary object detection (OVD). Unlike…
Graph-based Topology Reasoning for Driving Scenes
Tianyu Li, Li Chen, Huijie Wang +10
Understanding the road genome is essential to realize autonomous driving. This highly intelligent problem contains two aspects - the connection relationship of lanes, and the assig…