most citedAlignMamba: Enhancing Multimodal Mamba with Local and Global Cross-modal Alignment

1 citations · 1 across the 6 of their papers we have counts for

collaborators

6 papers

cs.LG2025

Bolster Hallucination Detection via Prompt-Guided Data Augmentation

Wenyun Li, Zheng Zhang, Dongmei Jiang +1

Large language models (LLMs) have garnered significant interest in AI community. Despite their impressive generation capabilities, they have been found to produce misleading or fab…

cs.CV2025

DS-Det: Single-Query Paradigm and Attention Disentangled Learning for Flexible Object Detection

Guiping Cao, Xiangyuan Lan, Wenjian Huang +3

Popular transformer detectors have achieved promising performance through query-based learning using attention mechanisms. However, the roles of existing decoder query types (e.g.,…

cs.CV2025

Cross-DINO: Cross the Deep MLP and Transformer for Small Object Detection

Guiping Cao, Wenjian Huang, Xiangyuan Lan +3

Small Object Detection (SOD) poses significant challenges due to limited information and the model's low class prediction score. While Transformer-based detectors have shown promis…

cs.CV2025

Open-Det: An Efficient Learning Framework for Open-Ended Detection

Guiping Cao, Tao Wang, Wenjian Huang +3

Open-Ended object Detection (OED) is a novel and challenging task that detects objects and generates their category names in a free-form manner, without requiring additional vocabu…

cs.CV2024

Transferable Adversarial Face Attack with Text Controlled Attribute

Wenyun Li, Zheng Zhang, Xiangyuan Lan +1

Traditional adversarial attacks typically produce adversarial examples under norm-constrained conditions, whereas unrestricted adversarial examples are free-form with semantically…

cs.CV20241 cited

AlignMamba: Enhancing Multimodal Mamba with Local and Global Cross-modal Alignment

Yan Li, Yifei Xing, Xiangyuan Lan +3

Cross-modal alignment is crucial for multimodal representation fusion due to the inherent heterogeneity between modalities. While Transformer-based methods have shown promising res…