collaborators

5 papers

cs.AI2026

OMG-Agent: Toward Robust Missing Modality Generation with Decoupled Coarse-to-Fine Agentic Workflows

Ruiting Dai, Zheyu Wang, Haoyu Yang +6

Data incompleteness severely impedes the reliability of multimodal systems. Existing reconstruction methods face distinct bottlenecks: conventional parametric/generative models are…

cs.SD2026

Do Models Hear Like Us? Probing the Representational Alignment of Audio LLMs and Naturalistic EEG

Haoyun Yang, Xin Xiao, Jiang Zhong +5

Audio Large Language Models (Audio LLMs) have demonstrated strong capabilities in integrating speech perception with language understanding. However, whether their internal represe…

cs.CV2026

LaViT: Aligning Latent Visual Thoughts for Multi-modal Reasoning

Linquan Wu, Tianxiang Jiang, Yifei Dong +6

Current multimodal latent reasoning often relies on external supervision (e.g., auxiliary images), ignoring intrinsic visual attention dynamics. In this work, we identify a critica…

cs.CV2025

FaceSleuth-R: Adaptive Orientation-Aware Attention for Robust Micro-Expression Recognition

Linquan Wu, Tianxiang Jiang, Haoyu Yang +5

Micro-expression recognition (MER) has achieved impressive accuracy in controlled laboratory settings. However, its real-world applicability faces a significant generalization clif…

cs.CL2025

SuperMerge: An Approach For Gradient-Based Model Merging

Haoyu Yang, Zheng Zhang, Saket Sathe

Large language models, such as ChatGPT, Claude, or LLaMA, are gigantic, monolithic, and possess the superpower to simultaneously support thousands of tasks. However, high-throughpu…