1 citations · 1 across the 3 of their papers we have counts for
3 papers
cs.CV2026
Reducing Object Hallucination in LVLMs via Emphasizing Image-negative Tokens
Meng Shen, Minghao Wu, Deepu Rajan
Object hallucination is a significant challenge that hinders the application of large vision-language models (LVLMs) in practice. We hypothesize that one possible origin of halluci…
cs.MM2024
Enhancing Modality Representation and Alignment for Multimodal Cold-start Active Learning
Meng Shen, Yake Wei, Jianxiong Yin +3
Training multimodal models requires a large amount of labeled data. Active learning (AL) aim to reduce labeling costs. Most AL methods employ warm-start approaches, which rely on s…
cs.MM2024★ 1 cited
Cross-Modality and Within-Modality Regularization for Audio-Visual DeepFake Detection
Heqing Zou, Meng Shen, Yuchen Hu +3
Audio-visual deepfake detection scrutinizes manipulations in public video using complementary multimodal cues. Current methods, which train on fused multimodal data for multimodal…