4 papers
What Guides the Agent? Adjudicating Unauthorized Behavior via Localizing Behavior-Guiding Instructions
Yichao Gao, Yumo Zhang, Yunhao Yao +4
LLM agents integrated with external resources gain complex task capabilities, yet the unified natural-language context channel makes them vulnerable to injection attacks: untrusted…
Coverage-Driven Adaptive Keyframe Selection for Video Understanding
Junyang Zhang, Puhan Luo, Chen Tang +2
Recent advances in large vision-language models (LVLMs) have enabled long-video understanding and analysis. However, processing the large number of frames in a video incurs substan…
ComPrivDet: Efficient Privacy Object Detection in Compressed Domains Through Inference Reuse
Yunhao Yao, Zhiqiang Wang, Ruiqi Li +3
As the Internet of Things (IoT) becomes deeply embedded in daily life, users are increasingly concerned about privacy leakage, especially from video data. Since frame-by-frame prot…
A-VL: Adaptive Attention for Large Vision-Language Models
Junyang Zhang, Mu Yuan, Ruiguang Zhong +5
The Large Vision-Language Model (LVLM) integrates computer vision and natural language processing techniques, offering substantial application potential. However, these models dema…