4 papers
Toward Natural and Companionable Virtual Agents via Cross-Temporal Emotional Modeling
Feier Qin, Xiao Li, Yi Zheng +5
Recent advances in foundation models have enabled conversational agents that aim for sustained companionship rather than mere task completion. Yet most still remain unable to suppo…
One-Step Diffusion-Based Image Compression with Semantic Distillation
Naifu Xue, Zhaoyang Jia, Jiahao Li +3
While recent diffusion-based generative image codecs have shown impressive performance, their iterative sampling process introduces unpleasing latency. In this work, we revisit the…
Universal Speech Token Learning via Low-Bitrate Neural Codec and Pretrained Representations
Xue Jiang, Xiulian Peng, Yuan Zhang +1
Current large speech language models are mainly based on semantic tokens from discretization of self-supervised learned representations and acoustic tokens from a neural codec, fol…
Generative Latent Coding for Ultra-Low Bitrate Image and Video Compression
Linfeng Qi, Zhaoyang Jia, Jiahao Li +3
Most existing approaches for image and video compression perform transform coding in the pixel space to reduce redundancy. However, due to the misalignment between the pixel-space…