4 papers
SignalLLM: A General-Purpose LLM Agent Framework for Automated Signal Processing
Junlong Ke, Qiying Hu, Shenghai Yuan +2
Modern signal processing (SP) pipelines, whether model-based or data-driven, often constrained by complex and fragmented workflow, rely heavily on expert knowledge and manual engin…
Semantic Surgery: Zero-Shot Concept Erasure in Diffusion Models
Lexiang Xiong, Chengyu Liu, Jingwen Ye +2
Concept erasure in text-to-image diffusion models is crucial for mitigating harmful content, yet existing methods often compromise generative quality. We introduce Semantic Surgery…
Interactive Test-Time Adaptation with Reliable Spatial-Temporal Voxels for Multi-Modal Segmentation
Haozhi Cao, Yuecong Xu, Pengyu Yin +4
Multi-modal test-time adaptation (MM-TTA) adapts models to an unlabeled target domain by leveraging the complementary multi-modal inputs in an online manner. While previous MM-TTA…
Minute-Long Videos with Dual Parallelisms
Zeqing Wang, Bowen Zheng, Xingyi Yang +3
Diffusion Transformer (DiT)-based video diffusion models generate high-quality videos at scale but incur prohibitive processing latency and memory costs for long videos. To address…