2 papers
cs.SD2026
FlashLabs Chroma 1.0: A Real-Time End-to-End Spoken Dialogue Model with Personalized Voice Cloning
Tanyu Chen, Tairan Chen, Kai Shen +4
Recent end-to-end spoken dialogue systems leverage speech tokenizers and neural audio codecs to enable LLMs to operate directly on discrete speech representations. However, these m…
cs.CV2024
Prompt-Guided Environmentally Consistent Adversarial Patch
Chaoqun Li, Huanqian Yan, Lifeng Zhou +3
Adversarial attacks in the physical world pose a significant threat to the security of vision-based systems, such as facial recognition and autonomous driving. Existing adversarial…