9 papers
Adaptive Forensic Feature Refinement via Intrinsic Importance Perception
Jiazhen Yang, Junjun Zheng, Kejia Chen +5
With the rapid development of generative models and multimodal content editing technologies, the key challenge faced by synthetic image detection (SID) lies in cross-distribution g…
Instinct vs. Reflection: Unifying Token and Verbalized Confidence in Multimodal Large Models
Yunkai Dang, Yifan Jiang, Yizhu Jiang +3
Multimodal Large Language Models (MLLMs) have demonstrated exceptional capabilities in various perception and reasoning tasks. Despite this success, ensuring their reliability in p…
Understanding and Enforcing Weight Disentanglement in Task Arithmetic
Shangge Liu, Yuehan Yin, Lei Wang +5
Task arithmetic provides an efficient, training-free way to edit pre-trained models, yet lacks a fundamental theoretical explanation for its success. The existing concept of ``weig…
UHR-BAT: Budget-Aware Token Compression Vision-Language model for Ultra-High-Resolution Remote Sensing
Yunkai Dang, Minxin Dai, Yuekun Yang +4
Ultra-high-resolution (UHR) remote sensing imagery couples kilometer-scale context with query-critical evidence that may occupy only a few pixels. Such vast spatial scale leads to…
CLASP: Class-Adaptive Layer Fusion and Dual-Stage Pruning for Multimodal Large Language Models
Yunkai Dang, Yizhu Jiang, Yifan Jiang +4
Multimodal Large Language Models (MLLMs) suffer from substantial computational overhead due to the high redundancy in visual token sequences. Existing approaches typically address…
Distractor-free Generalizable 3D Gaussian Splatting
Yanqi Bao, Jing Liao, Jing Huo +1
We present DGGS, a novel framework that addresses the previously unexplored challenge: (3DGS). It mitigates 3D incons…