From the 1 of 24 linked papers with an AI index.
24 papers
PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs
Zhongjie Ba, Shengwang Xu, Peng Cheng +4
Embodied intelligence and world models require video understanding systems to go beyond recognizing objects and actions and develop an understanding of physical regularities. Howev…
Inverting the Hidden: Unveiling Multimodal Privacy Leakage in Collaborative LVLM Inference
Shuaifan Jin, Zhibo Wang, Qiyuan Wang +5
Collaborative inference deploys Large Vision-Language Models (LVLMs) by partitioning computation between edge devices and the cloud. While withholding raw inputs supposedly ensures…
From Role Prompt to Infinite Thinking: Exploiting Persona Conditioning for Inference Cost Attacks in LLMs
Zhiyi Mou, Wangze Ni, Tianfang Xiao +6
LLMs are increasingly deployed in real-world applications, making inference efficiency and service reliability critical concerns due to their substantial computational costs. Howev…
Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models
Xuanyi Hao, Zuoyuan Zhang, Zhibo Wang +4
The paper introduces a plug‑and‑play, attention‑free token reduction module for vision‑language models that selects informative and diverse visual tokens using an entropy‑based imp…
LoRA-Key: User-Centric LoRA Watermarking for Text-to-Image Diffusion Models
Yaopeng Wang, Qingliang Wang, Zhibo Wang +5
Low-Rank Adaptation (LoRA) has become a widely used mechanism for customizing text-to-image diffusion models, enabling lightweight modules that are shared, reused, and commercializ…
RouteScan: A Non-Intrusive Approach to Auditing MoE LLMs Safety via Expert Routing Telemetry
Bo Lv, Zhiheng Xu, KeDong Xiu +4
As Mixture-of-Experts (MoE) architectures are increasingly adopted for scaling Large Language Models (LLMs), safety auditing becomes necessary to verify whether these models produc…