11 papers
Tackling Multimodal Learning Challenges with Mixture-of-Expert: A Survey
Liangwei Nathan Zheng, Wei Emma Zhang, Olaf Maennel +2
Mixture-of-Experts (MoE) presents a naturally compatible and scalable framework for multimodal learning, demonstrating strong adaptability across diverse modalities and tasks. Desp…
Probing Routing-Conditional Calibration in Attention-Residual Transformers
Wenhao Liang, Lin Yue, Wei Emma Zhang +4
Post-hoc calibration is usually evaluated as a function of logits or softmax confidence alone, even as routing-augmented architectures increasingly accompany predictions with sampl…
Rethinking Gating Mechanism in Sparse MoE: Handling Arbitrary Modality Inputs with Confidence-Guided Gate
Liangwei Nathan Zheng, Wei Emma Zhang, Mingyu Guo +2
Effectively managing missing modalities is a fundamental challenge in real-world multimodal learning scenarios, where data incompleteness often results from systematic collection e…
Test-Time Attention Purification for Backdoored Large Vision Language Models
Zhifang Zhang, Bojun Yang, Shuo He +5
Despite the strong multimodal performance, large vision-language models (LVLMs) are vulnerable during fine-tuning to backdoor attacks, where adversaries insert trigger-embedded sam…
Calibration Attention: Learning Reliability-Aware Representations for Vision Transformers
Wenhao Liang, Wei Emma Zhang, Lin Yue +4
Most calibration methods operate at the logit level, implicitly assuming that miscalibration can be corrected without changing the underlying representation. We challenge this assu…
Lifting Manifolds to Mitigate Pseudo-Alignment in LLM4TS
Liangwei Nathan Zheng, Wenhao Liang, Wei Emma Zhang +3
Pseudo-Alignment is a pervasive challenge in many large language models for time series (LLM4TS) models, often causing them to underperform compared to linear models or randomly in…