10 papers
Intern-S1-Pro: Scientific Multimodal Foundation Model at Trillion Scale
Yicheng Zou, Dongsheng Zhu, Lin Zhu +174
We introduce Intern-S1-Pro, the first one-trillion-parameter scientific multimodal foundation model. Scaling to this unprecedented size, the model delivers a comprehensive enhancem…
STEP: Scientific Time-Series Encoder Pretraining via Cross-Domain Distillation
Chen Zhang, Liwei Liu, Jun Tao +6
Scientific time series are central to scientific AI but are typically sparse, highly heterogeneous, and limited in scale, making unified representation learning particularly challe…
SemanticVocoder: Bridging Audio Generation and Audio Understanding via Semantic Latents
Zeyu Xie, Chenxing Li, Qiao Jin +6
Recent audio generation models typically rely on Variational Autoencoders (VAEs) and perform generation within the VAE latent space. Although VAEs excel at compression and reconstr…
HoliAntiSpoof: Audio LLM for Holistic Speech Anti-Spoofing
Xuenan Xu, Yiming Ren, Liwei Liu +5
Recent advances in speech synthesis and editing have made speech spoofing increasingly challenging. However, most existing methods treat spoofing as binary classification, overlook…
MMEDIT: A Unified Framework for Multi-Type Audio Editing via Audio Language Model
Ye Tao, Wen Wu, Chao Zhang +3
Text-guided audio editing aims to modify specific acoustic events while strictly preserving non-target content. Despite recent progress, existing approaches remain fundamentally li…
HIPPO: Accelerating Video Large Language Models Inference via Holistic-aware Parallel Speculative Decoding
Qitan Lv, Tianyu Liu, Wen Wu +4
Speculative decoding (SD) has emerged as a promising approach to accelerate LLM inference without sacrificing output quality. Existing SD methods tailored for video-LLMs primarily…