activity
20242026
collaborators

16 papers

cs.CV2026

Modality Gap-Driven Subspace Alignment Training Paradigm For Multimodal Large Language Models

Xiaomin Yu, Yi Xin, Yuhui Zhang +12

Despite the success of multimodal contrastive learning in aligning visual and linguistic representations, a persistent geometric anomaly, the Modality Gap, remains: embeddings of d…

cs.CV2026

ReactionMamba: Generating Short & Long Human Reaction Sequences

Hajra Anwar Beg, Baptiste Chopin, Hao Tang +1

We present ReactionMamba, a novel framework for generating long 3D human reaction motions. Reaction-Mamba integrates a motion VAE for efficient motion encoding with Mamba-based sta…

cs.MM2025

An Evaluation of Interleaved Instruction Tuning on Semantic Reasoning Performance in an Audio MLLM

Jiawei Liu, Enis Berk Çoban, Zarina Schevchenko +4

Standard training for Multi-modal Large Language Models (MLLMs) involves concatenating non-textual information, like vision or audio, with a text prompt. This approach may not enco…

cs.CV2025

Multimodal Alignment and Fusion: A Survey

Songtao Li, Hao Tang

This survey provides a comprehensive overview of recent advances in multimodal alignment and fusion within the field of machine learning, driven by the increasing availability and…

cs.CV2025

Resolving Task Objective Conflicts in Unified Model via Task-Aware Mixture-of-Experts

Jiaxing Zhang, Hao Tang

Unified multimodal large language models (MLLMs) based on end-to-end autoregressive (AR) transformers effectively integrate both understanding and generation tasks within a single…

cs.CV2025

PolSAM: Polarimetric Scattering Mechanism Informed Segment Anything Model

Yuqing Wang, Zhongling Huang, Shuxin Yang +4

PolSAR data presents unique challenges due to its rich and complex characteristics. Existing data representations, such as complex-valued data, polarimetric features, and amplitude…