3 papers
cs.CV2026
PEAR: Pixel-aligned Expressive humAn mesh Recovery
Jiahao Wu, Yunfei Liu, Lijian Lin +4
Reconstructing detailed 3D human meshes from a single in-the-wild image remains a fundamental challenge in computer vision. Existing SMPLX-based methods often suffer from slow infe…
cs.SD2025
MelTok: 2D Tokenization for Single-Codebook Audio Compression
Jingyi Li, Zhiyuan Zhao, Zhisheng Zhang +6
Large Audio Language Models (LALMs) have emerged with strong performance across diverse audio understanding tasks and can be further enhanced by neural audio codecs. Transitioning…
cs.SD2025
Training-Free Multi-Step Audio Source Separation
Yongyi Zang, Jingyi Li, Qiuqiang Kong
Audio source separation aims to separate a mixture into target sources. Previous audio source separation systems usually conduct one-step inference, which does not fully explore th…