collaborators

9 papers

cs.SD2026

Audio Interaction Model

Zhifei Xie, Zihang Liu, Ze An +8

Audio is continuous and interactive, yet most Large Audio Language Models (LALMs) remain offline and streaming systems usually specialize in ASR or spoken dialogue. We formalize th…

cs.CV2025

UniVST: A Unified Framework for Training-free Localized Video Style Transfer

Quanjian Song, Mingbao Lin, Wengyi Zhan +3

This paper presents UniVST, a unified framework for localized video style transfer based on diffusion models. It operates without the need for training, offering a distinct advanta…

cs.SD2025

Audio-Reasoner: Improving Reasoning Capability in Large Audio Language Models

Zhifei Xie, Mingbao Lin, Zihang Liu +3

Recent advancements in multimodal reasoning have largely overlooked the audio modality. We introduce Audio-Reasoner, a large-scale audio language model for deep reasoning in audio…

cs.CL2025

UIO-LLMs: Unbiased Incremental Optimization for Long-Context LLMs

Wenhao Li, Mingbao Lin, Yunshan Zhong +2

Managing long texts is challenging for large language models (LLMs) due to limited context window sizes. This study introduces UIO-LLMs, an unbiased incremental optimization approa…

cs.CV2025

EasyInv: Toward Fast and Better DDIM Inversion

Ziyue Zhang, Mingbao Lin, Shuicheng Yan +1

This paper introduces EasyInv, an easy yet novel approach that significantly advances the field of DDIM Inversion by addressing the inherent inefficiencies and performance limitati…

cs.CV2025

DiffusionTrend: A Minimalist Approach to Virtual Fashion Try-On

Wengyi Zhan, Mingbao Lin, Shuicheng Yan +1

We introduce DiffusionTrend for virtual fashion try-on, which forgoes the need for retraining diffusion models. Using advanced diffusion models, DiffusionTrend harnesses latent inf…