24 papers
ARIES-Mission2: A Zero-Shot Vision-Language-Action Framework for Fast Large-Scale Aerial Mission Generation
Junhao Wei, Yanxiao Li, Haochen Li +8
Multimodal Large Language Models (MLLMs) have shown strong semantic understanding capabilities, but their direct use in low-altitude Unmanned Aerial Vehicle (UAV) mission generatio…
Remember-R1: Mitigating Long-Context Visual Forgetting through Reinforcement Learning
Jianmin Chen, Jiaqi Tang, Wei Wei +9
Multimodal large language models (MLLMs) increasingly rely on long chain-of-thought reasoning for complex tasks. However, as reasoning sequences lengthen, models may gradually rely…
Unveiling the Unknown: Open Vocabulary Object Detection with Scene Graphs
Yi Chen, Yinghao Lu, Zhehao Li +4
Open-vocabulary object detection seeks to identify novel object categories that were not part of the training data. Many knowledge distillation-based approaches have shown promisin…
HoloFair: Unified T2I Fairness Evaluation and Fair-GRPO Debiasing
Ruyi Chen, Lu Zhou, Xiaogang Xu +3
Text-to-Image (T2I) models have made significant strides in visual realism and semantic consistency, yet they often perpetuate and amplify societal biases. Existing evaluation meth…
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments
Chiyu Zhang, Huiqin Yang, Bendong Jiang +8
The rapid proliferation of LLM-based autonomous agents in real operating system environments introduces a new category of safety risk beyond content safety: behavior jailbreak, whe…
Restoration-Oriented Video Frame Interpolation with Region-Distinguishable Priors from SAM
Yan Han, Xiaogang Xu, Yingqi Lin +3
In existing restoration-oriented Video Frame Interpolation (VFI) approaches, the motion estimation between neighboring frames plays a crucial role. However, the estimation accuracy…