activity
20242026
collaborators

9 papers

cs.CV2026

Adaptive Forensic Feature Refinement via Intrinsic Importance Perception

Jiazhen Yang, Junjun Zheng, Kejia Chen +5

With the rapid development of generative models and multimodal content editing technologies, the key challenge faced by synthetic image detection (SID) lies in cross-distribution g…

cs.CV2026

Instinct vs. Reflection: Unifying Token and Verbalized Confidence in Multimodal Large Models

Yunkai Dang, Yifan Jiang, Yizhu Jiang +3

Multimodal Large Language Models (MLLMs) have demonstrated exceptional capabilities in various perception and reasoning tasks. Despite this success, ensuring their reliability in p…

cs.AI2026

Understanding and Enforcing Weight Disentanglement in Task Arithmetic

Shangge Liu, Yuehan Yin, Lei Wang +5

Task arithmetic provides an efficient, training-free way to edit pre-trained models, yet lacks a fundamental theoretical explanation for its success. The existing concept of ``weig…

cs.CV2026

UHR-BAT: Budget-Aware Token Compression Vision-Language model for Ultra-High-Resolution Remote Sensing

Yunkai Dang, Minxin Dai, Yuekun Yang +4

Ultra-high-resolution (UHR) remote sensing imagery couples kilometer-scale context with query-critical evidence that may occupy only a few pixels. Such vast spatial scale leads to…

cs.CV2026

CLASP: Class-Adaptive Layer Fusion and Dual-Stage Pruning for Multimodal Large Language Models

Yunkai Dang, Yizhu Jiang, Yifan Jiang +4

Multimodal Large Language Models (MLLMs) suffer from substantial computational overhead due to the high redundancy in visual token sequences. Existing approaches typically address…

cs.CV2026

Distractor-free Generalizable 3D Gaussian Splatting

Yanqi Bao, Jing Liao, Jing Huo +1

We present DGGS, a novel framework that addresses the previously unexplored challenge: (3DGS). It mitigates 3D incons…