6 papers · 1 filter
MotionCraft: Latent World Modeling with Sparse Attention for Visual Upscaling
Rong Fu, Chunlei Meng, Yangchen Zeng +9
Video super-resolution (VSR) aims to recover high-fidelity high-resolution videos from low-resolution inputs and is central to applications ranging from mobile capture to streaming…
Do Vision Models Truly Forget? New Findings from Representation-Level Certification of Visual Unlearning in Vertical Federated Learning
Zhenyu Yu, Yangchen Zeng, Chunlei Meng +2
Machine unlearning in Vertical Federated Learning (VFL) has attracted growing interest, yet existing methods certify forgetting solely using output-level metrics. We challenge thes…
SpaMEM: Benchmarking Dynamic Spatial Reasoning via Perception-Memory Integration in Embodied Environments
Chih-Ting Liao, Xi Xiao, Chunlei Meng +6
Multimodal large language models (MLLMs) have advanced static visual--spatial reasoning, yet they often fail to preserve long-horizon spatial coherence in embodied settings where b…
CoDA: Exploring Chain-of-Distribution Attacks and Post-Hoc Token-Space Repair for Medical Vision-Language Models
Xiang Chen, Fangfang Yang, Chunlei Meng +6
Medical vision--language models (MVLMs) are increasingly used as perceptual backbones in radiology pipelines and as the visual front end of multimodal assistants, yet their reliabi…
DIVER: Dynamic Iterative Visual Evidence Reasoning for Multimodal Fake News Detection
Weilin Zhou, Zonghao Ying, Chunlei Meng +6
Multimodal fake news detection is crucial for mitigating adversarial misinformation. Existing methods, relying on static fusion or LLMs, face computational redundancy and hallucina…
Adversarial Robustness for Unified Multi-Modal Encoders via Efficient Calibration
Chih-Ting Liao, Zhangquan Chen, Chunlei Meng +3
Recent unified multi-modal encoders align a wide range of modalities into a shared representation space, enabling diverse cross-modal tasks. Despite their impressive capabilities,…