activity
20242026
collaborators

5 papers

cs.CV2026

Nüwa: Mending the Spatial Integrity Torn by VLM Token Pruning

Yihong Huang, Fei Ma, Yihua Shao +4

Vision token pruning has proven to be an effective acceleration technique for the efficient Vision Language Model (VLM). However, existing pruning methods demonstrate excellent per…

cs.RO2025

H-GAR: A Hierarchical Interaction Framework via Goal-Driven Observation-Action Refinement for Robotic Manipulation

Yijie Zhu, Rui Shao, Ziyang Liu +4

Unified video and action prediction models hold great potential for robotic manipulation, as future observations offer contextual cues for planning, while actions reveal how intera…

cs.CV2025

SAC-MIL: Spatial-Aware Correlated Multiple Instance Learning for Histopathology Whole Slide Image Classification

Yu Bai, Zitong Yu, Haowen Tian +11

We propose Spatial-Aware Correlated Multiple Instance Learning (SAC-MIL) for performing WSI classification. SAC-MIL consists of a positional encoding module to encode position info…

cs.CV2025

Multimodal Fake News Detection: MFND Dataset and Shallow-Deep Multitask Learning

Ye Zhu, Yunan Wang, Zitong Yu

Multimodal news contains a wealth of information and is easily affected by deepfake modeling attacks. To combat the latest image and text generation methods, we present a new Multi…

cs.CV2024

EPE-P: Evidence-based Parameter-efficient Prompting for Multimodal Learning with Missing Modalities

Zhe Chen, Xun Lin, Yawen Cui +1

Missing modalities are a common challenge in real-world multimodal learning scenarios, occurring during both training and testing. Existing methods for managing missing modalities…