8 papers
VSD2M: A Large-scale Vision-language Sticker Dataset for Multi-frame Animated Sticker Generation
Zhiqiang Yuan, Jiapei Zhang, Ying Deng +3
As a common form of communication in social media,stickers win users' love in the internet scenarios, for their ability to convey emotions in a vivid, cute, and interesting way. Pe…
Exploiting Prefix-Tree in Structured Output Interfaces for Enhancing Jailbreak Attacking
Yanzeng Li, Yunfan Xiong, Jialun Zhong +3
The rise of Large Language Models (LLMs) has led to significant applications but also introduced serious security threats, particularly from jailbreak attacks that manipulate outpu…
CAB: Comprehensive Attention Benchmarking on Long Sequence Modeling
Jun Zhang, Shuyang Jiang, Jiangtao Feng +2
Transformer has achieved remarkable success in language, image, and speech processing. Recently, various efficient attention architectures have been proposed to improve transformer…
SuperFusion: Multilevel LiDAR-Camera Fusion for Long-Range HD Map Generation
Hao Dong, Weihao Gu, Xianjing Zhang +5
High-definition (HD) semantic map generation of the environment is an essential component of autonomous driving. Existing methods have achieved good performance in this task by fus…
Leveraging Diverse Semantic-based Audio Pretrained Models for Singing Voice Conversion
Xueyao Zhang, Zihao Fang, Yicheng Gu +5
Singing Voice Conversion (SVC) is a technique that enables any singer to perform any song. To achieve this, it is essential to obtain speaker-agnostic representations from the sour…
MedDiT: A Knowledge-Controlled Diffusion Transformer Framework for Dynamic Medical Image Generation in Virtual Simulated Patient
Yanzeng Li, Cheng Zeng, Jinchao Zhang +2
Medical education relies heavily on Simulated Patients (SPs) to provide a safe environment for students to practice clinical skills, including medical image analysis. However, the…