4 papers
Diffusion Model as a Generalist Segmentation Learner
Haoxiao Wang, Antao Xiang, Haiyang Sun +8
Diffusion models are primarily trained for image synthesis, yet their denoising trajectories encode rich, spatially aligned visual priors. In this paper, we demonstrate that these…
WavAlign: Enhancing Intelligence and Expressiveness in Spoken Dialogue Models via Adaptive Hybrid Post-Training
Yifu Chen, Shengpeng Ji, Qian Chen +9
End-to-end spoken dialogue models have garnered significant attention because they offer a higher potential ceiling in expressiveness and perceptual ability than cascaded systems.…
TransDiff: Diffusion-Based Method for Manipulating Transparent Objects Using a Single RGB-D Image
Haoxiao Wang, Kaichen Zhou, Binrui Gu +6
Manipulating transparent objects presents significant challenges due to the complexities introduced by their reflection and refraction properties, which considerably hinder the acc…
WavRAG: Audio-Integrated Retrieval Augmented Generation for Spoken Dialogue Models
Yifu Chen, Shengpeng Ji, Haoxiao Wang +5
Retrieval Augmented Generation (RAG) has gained widespread adoption owing to its capacity to empower large language models (LLMs) to integrate external knowledge. However, existing…