8 papers
Nüwa: Mending the Spatial Integrity Torn by VLM Token Pruning
Yihong Huang, Fei Ma, Yihua Shao +4
Vision token pruning has proven to be an effective acceleration technique for the efficient Vision Language Model (VLM). However, existing pruning methods demonstrate excellent per…
H-GAR: A Hierarchical Interaction Framework via Goal-Driven Observation-Action Refinement for Robotic Manipulation
Yijie Zhu, Rui Shao, Ziyang Liu +4
Unified video and action prediction models hold great potential for robotic manipulation, as future observations offer contextual cues for planning, while actions reveal how intera…
SAC-MIL: Spatial-Aware Correlated Multiple Instance Learning for Histopathology Whole Slide Image Classification
Yu Bai, Zitong Yu, Haowen Tian +11
We propose Spatial-Aware Correlated Multiple Instance Learning (SAC-MIL) for performing WSI classification. SAC-MIL consists of a positional encoding module to encode position info…
Multimodal Fake News Detection: MFND Dataset and Shallow-Deep Multitask Learning
Ye Zhu, Yunan Wang, Zitong Yu
Multimodal news contains a wealth of information and is easily affected by deepfake modeling attacks. To combat the latest image and text generation methods, we present a new Multi…
Dynamic Analysis and Adaptive Discriminator for Fake News Detection
Xinqi Su, Zitong Yu, Yawen Cui +7
In current web environment, fake news spreads rapidly across online social networks, posing serious threats to society. Existing multimodal fake news detection methods can generall…
FusionMamba: Dynamic Feature Enhancement for Multimodal Image Fusion with Mamba
Xinyu Xie, Yawen Cui, Tao Tan +2
Multimodal image fusion aims to integrate information from different imaging techniques to produce a comprehensive, detail-rich single image for downstream vision tasks. Existing m…