3 papers
cs.AI2026
Soul Computing: A Theoretical Framework and Technical Architecture for Intelligent Agents with Independent Consciousness
Jinshan Zhang, Xishi Zhou, Qiu Peng +1
Breakthroughs in large language models and multimodal generation technologies have propelled the digital reconstruction of human mental traits, emotional patterns, and long-term me…
cs.SD2026
AST: Adaptive, Seamless, and Training-Free Precise Speech Editing
Sihan Lv, Yechen Jin, Zhen Li +5
Text-based speech editing aims to modify specific segments while preserving speaker identity and acoustic context. Current approaches generally involve either expensive task-specif…
cs.CV2026
OmniEdit: A Training-free framework for Lip Synchronization and Audio-Visual Editing
Lixiang Lin, Siyuan Jin, Jinshan Zhang
Lip synchronization and audio-visual editing have emerged as fundamental challenges in multimodal learning, underpinning a wide range of applications, including film production, vi…