From the 1 of 8 linked papers with an AI index.
8 papers
QuantWAMs: Calibrating at the Right Granularity for World Action Models
Jiacheng Zhou, Jinfan Lv, Ruixuan Li +4
The paper proposes QuantWAMs, a post‑training quantization framework that tailors quantization decisions to the structure, rollout distribution, and task objectives of World Action…
VLA-Hijack: A Transferable Patch Attack against Vision-Language-Action Models via Visual Proprioception Hijacking
Jiyuan Fu, Kaixun Jiang, Jingkai Jia +7
While Vision-Language-Action (VLA) models have emerged as powerful generalist policies, their severe vulnerability to adversarial patches significantly hinders their deployment in…
Unified Multimodal Visual Tracking with Dual Mixture-of-Experts
Lingyi Hong, Jinglun Li, Xinyu Zhou +6
Multimodal visual object tracking can be divided into to several kinds of tasks (e.g. RGB and RGB+X tracking), based on the input modality. Existing methods often train separate mo…
LingoLoop Attack: Trapping MLLMs via Linguistic Context and State Entrapment into Endless Loops
Jiyuan Fu, Kaixun Jiang, Lingyi Hong +5
Multimodal Large Language Models (MLLMs) have shown great promise but require substantial computational resources during inference. Attackers can exploit this by inducing excessive…
GenAgent: Scaling Text-to-Image Generation via Agentic Multimodal Reasoning
Kaixun Jiang, Yuzheng Wang, Junjie Zhou +6
We introduce GenAgent, unifying visual understanding and generation through an agentic multimodal model. Unlike unified models that face expensive training costs and understanding-…
Dynamic Semantic-Aware Correlation Modeling for UAV Tracking
Xinyu Zhou, Tongxin Pan, Lingyi Hong +5
UAV tracking can be widely applied in scenarios such as disaster rescue, environmental monitoring, and logistics transportation. However, existing UAV tracking methods predominantl…