Publications (184)
Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing
Fengxiang Wang, Jiangnan Huang, Mingshuo Chen +8
Ultra-high-resolution (UHR) remote-sensing (RS) imagery provides fine-grained Earth-observation evidence over city-scale scenes, but poses a fundamental challenge for multimodal la…
Beyond the Last Layer: Multi-Layer Representation Fusion for Visual Tokenization
Xuanyu Zhu, Yan Bai, Yang Shi +4
Representation autoencoders that reuse frozen pretrained vision encoders as visual tokenizers have achieved strong reconstruction and generation quality. However, existing methods…
Architecture-Agnostic Feature Synergy for Universal Defense Against Heterogeneous Generative Threats
Bingxue Zhang, Yang Gao, Feida Zhu +2
Generative AI deployment poses unprecedented challenges to content safety and privacy. However, existing defense mechanisms are often tailored to specific architectures (e.g., Diff…
Connectivity-Preserving Consensus of Multi-Agent Systems with Bounded Actuation
Yuan Yang, Daniela Constantinescu, Yang Shi
This paper investigates the impact of bounded actuation on the connectivity-preserving consensus of two classes of multi-agent systems, with kinematic agents and with Euler- Lagran…
A Robust Distributed Model Predictive Control Framework for Consensus of Multi-Agent Systems with Input Constraints and Varying Delays
Henglai Wei, Changxin Liu, Yang Shi
This paper studies the consensus problem of general linear discrete-time multi-agent systems (MAS) with input constraints and bounded time-varying communication delays. We propose…
VTC-Bench: Evaluating Agentic Multimodal Models via Compositional Visual Tool Chaining
Xuanyu Zhu, Yuhao Dong, Rundong Wang +9
Recent advancements extend Multimodal Large Language Models (MLLMs) beyond standard visual question answering to utilizing external tools for advanced visual tasks. Despite this pr…