3 papers
cs.CV2026
LAST: The Last Query Token Guides Visual Token Pruning for Edge-Cloud Collaborative MLLM Inference
Feng Yang, Xinrui Ju, Keyang Zhang +6
Multimodal foundation models are reshaping edge-cloud visual intelligence from task-specific feature pipelines into token-based interfaces, where edge devices encode visual inputs…
cs.CV2025
Propose and Rectify: A Forensics-Driven MLLM Framework for Image Manipulation Localization
Keyang Zhang, Chenqi Kong, Hui Liu +3
The increasing sophistication of image manipulation techniques demands robust forensic solutions that can both reliably detect alterations and precisely localize tampered regions.…
eess.IV2024
Image Provenance Analysis via Graph Encoding with Vision Transformer
Keyang Zhang, Chenqi Kong, Shiqi Wang +2
Recent advances in AI-powered image editing tools have significantly lowered the barrier to image modification, raising pressing security concerns those related to spreading misinf…