collaborators

12 papers

eess.AS2026

SALMONN-2: Advancing General-Purpose Hearing Abilities with Self-Supervised Representations

Xiaoyu Yang, Xuenan Xu, Wenyi Yu +10

Recent audio large language models (ALLMs) are typically built upon audio encoders trained with large amounts of supervised data. Since self-supervised learning (SSL) audio encoder…

cs.CV2026

JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments

Zhan Liu, Changli Tang, Yuxin Wang +7

Current audio-visual large language models (AV-LLMs) are predominantly restricted to 2D perception, relying on RGB video and monaural audio. This design choice introduces a fundame…

q-bio.NC2026

One Brain, Omni Modalities: Towards Unified Non-Invasive Brain Decoding with Large Language Models

Changli Tang, Shurui Li, Junliang Wang +8

Deciphering brain function through non-invasive recordings requires synthesizing complementary high-frequency electromagnetic (EEG/MEG) and low-frequency metabolic (fMRI) signals.…

cs.CV2026

WAVE: Learning Unified & Versatile Audio-Visual Embeddings with Multimodal LLM

Changli Tang, Qinfan Xiao, Ke Mei +3

While embeddings from multimodal large language models (LLMs) excel as general-purpose representations, their application to dynamic modalities like audio and video remains underex…

cs.CV2026

D-ORCA: Dialogue-Centric Optimization for Robust Audio-Visual Captioning

Changli Tang, Tianyi Wang, Fengyun Rao +2

Spoken dialogue is a primary source of information in videos; therefore, accurately identifying who spoke what and when is essential for deep video understanding. We introduce D-OR…

cs.SD2026

OCR-Enhanced Multimodal ASR Can Read While Listening

Junli Chen, Changli Tang, Yixuan Li +2

Visual information, such as subtitles in a movie, often helps automatic speech recognition. In this paper, we propose Donut-Whisper, an audio-visual ASR model with dual encoder to…